Deployment and Main-Project Integration
This page is the handoff contract for integrating the mapper into the larger fax-automation service. The mapper is a terminology and safety service; it is not an OCR engine, fax transport, patient-matching service, or Elation client. Those responsibilities remain in the main project.
Production Boundary
fax/PDF
-> main project's OCR and row extraction
-> LabObservation adapter
-> one long-lived Mapper per worker
-> MappingResult
-> accepted-only ElationAdapter payload
-> Elation API
The mapper must receive the complete row context whenever it is available: name, value, unit, specimen, collection specimen, method, time context, scale, reference range, source laboratory, observation domain, LOINC class hint, panel context, and report context. Do not reduce a row to only its name if the OCR system extracted a unit or specimen. Unit and axis evidence are safety inputs, not optional decoration.
For a recognized report heading, also send panel_context_source,
panel_context_confidence, panel_loinc_hint when it is known from trusted
local configuration, and a stable report_section_id. An exact official
heading uses confidence 0.98. These fields let the mapper verify component
membership; they never allow a parent panel code to be emitted for an
individual result row. See Panel context.
source_laboratory is optional enrichment, not a required routing key. An
ordinary result must map through the release-derived universal axis path when
the laboratory is absent or appears for the first time. Use
collection_specimen for a draw/source (for example Blood, Venous) and use
specimen only when the report identifies the analytical material. LOINC
class and panel context are advisory retrieval/ranking hints; LAB_RESULT is
the hard boundary that limits output to laboratory result terms.
For a repaired OCR table row, preserve raw_name exactly and pass the separate
mapping_name, input_quality, repair_operations, repair_confidence, and
source_cell_references fields. mapping_name is accepted only for
deterministic_repair or gemini_verified_repair. If the cells are merged or
shifted, send input_structure_invalid and do not fabricate a name/value/unit
combination. The mapper will produce an actionable extraction abstention before
candidate retrieval.
For a typed integration, send the canonical specimen enums in addition to the raw OCR text:
LabObservation(
raw_name="Hemoglobin",
value="12.0",
unit="g/dL",
specimen="blood",
specimen_code="blood_unspecified",
collection_specimen="Blood, Venous",
collection_specimen_code="blood_venous",
)
An equivalent reconstructed row keeps its source cell visible:
LabObservation(
raw_name="Hemoglobin 02",
mapping_name="Hemoglobin",
input_quality="deterministic_repair",
repair_operations=("remove_repeated_layout_footnote",),
repair_confidence=1.0,
source_cell_references=("page-2:cell-11",),
value="13.8",
unit="g/dL",
observation_domain="LAB_RESULT",
)
Analytical values are unknown, blood_unspecified, whole_blood, serum,
plasma, serum_or_plasma, urine, cerebrospinal_fluid, saliva, stool,
sputum, amniotic_fluid, synovial_fluid, pleural_fluid,
peritoneal_fluid, body_fluid_unspecified, tissue, and other.
Collection values are unknown, blood_venous, blood_arterial,
blood_capillary, urine_clean_catch, urine_catheter, urine_24_hour, and
other. Preserve the original words for audit. Bld is a LOINC System-axis
value, not a UI specimen enum. Contradictory text and enum values are rejected.
Artifact Bundle
Build or provision one matched artifact bundle. Do not mix releases between these files. A worker needs the first group; the raw UMLS files are retained only in the protected build/audit location after the serving index is built:
assets/loinc/2.82/catalog_v5.sqlite3
assets/loinc/2.82/catalog_v5.manifest.json
assets/loinc/2.82/terms.faiss
assets/loinc/2.82/terms.json
assets/loinc/2.82/terms.code_vectors.npy
assets/loinc/2.82/terms.code_vector_codes.json
assets/loinc/2.82/shards/vector_shards.json # optional, benchmarked only
assets/umls/2026AA/serving/umls_loinc.sqlite3
assets/umls/2026AA/serving/serving_manifest.json
config/mapping_registry.json
config/axis_context_aliases.json
config/axis_safety_policy.json
config/lab_panel_default_specimens.csv
config/panel_context_aliases.2.82.json
config/panel_context_index.2.82.json
config/laboratory_vendors.json
config/ocr_corrections.json
config/unit_surface_rules.json
The registry path is an explicit immutable runtime input, not a mutable file
inside the worker image. The certified OCR-recovery bootstrap is
config/mapping_registry.review_cases.20260910.json; revision
review-cases-20260910T000000Z-r1 has parent igenex-20260828T000000Z, keeps
the parent's 899 mappings, adds two normal clinician-approved mappings, and
contains 41 exact force overrides. It may be promoted directly only if that is
also the main project's current active parent. If the active pointer is newer,
the registry publisher must rebase the append-only additions onto the newer
parent, validate the resulting active snapshot, and then move the pointer
atomically. Never copy this file over a newer registry or hand-edit the
force_exact section in a running deployment.
To prepare a candidate rebased on a downloaded newer active snapshot, run this in the package repository, then have the main-project publisher validate its checksum and atomically promote it through the normal active-pointer workflow:
$env:PYTHONPATH = "$PWD\src"
python tools\build_review_case_snapshot.py `
--base C:\private\registry\current-active.json `
--catalog assets\loinc\2.82\catalog_v5.sqlite3 `
--output config\mapping_registry.review_cases.rebased.json
The command creates a new append-only candidate; it never modifies the base file or moves an active Cloud Storage pointer.
A publish round compiled from stored review cases follows the same lineage
rule through tools/compile_clinician_approvals.py (see CLI): the
active snapshot the main project loads is the parent, the finalized file
carries registry_status: active only after the strict full and fast-path
replays passed, and the manifest beside it records the parent version, the
replay report hashes and the registry hash. The 2026-09-18 round is
config/mapping_registry.review_cases.20260918.json
(registry-20260918T140454-284e8050dbb6, parent
review-cases-20260910T000000Z, 919 mappings, 41 force overrides). The
operator promotes it by syncing the file and manifest to the asset bucket and
pointing the service's registry path at it; /health then shows the new
version, status and hash.
The main project's publisher then descends from that file: the active
pointer in its data bucket (registry_publications/ACTIVE) named
registry-20260918T191655-f861427aad03 on 2026-09-22 (921 mappings, 41
force overrides; two clinician decisions made in the application on
2026-09-18 on top of the config snapshot). That snapshot is kept in the
repository as config/mapping_registry.published.20260918T191655.json (+
manifest; application case identifiers in its reason strings are redacted)
so the lineage of the next round is on disk: a round compiled here parents
on the active descendant, never on the older config file.
The 2026-09-22 round is config/mapping_registry.review_cases.20260922.json
(+ manifest; parent registry-20260918T191655-f861427aad03), compiled with
tools/compile_clinician_approvals.py --facts from the PHI-free
evaluation/clinical_review/publish_round_20260922.facts.json and
publish_round_20260922.json under Dr Tro: % SATURATION -> 2502-3 once
per iron panel (required_panel), PSA, % FREE -> 12841-3,
RDW Standard Deviation -> 21000-5, Nucleated RBCs, Automated -> 58413-6,
T4 (THYROXINE), TOTAL -> 3026-2 (flagged for Dr Tro) and the practice's
order-set surfaces for codes already approved. The operator publishes it
the way the application does: copy the file and its manifest into the data
bucket's registry_publications/ and point ACTIVE at it; the service pin
may stay on the config file because the application loads the descendant.
The same sync carries config/laboratory_vendors.json and the 2.82.5 panel
index (2.82.4 added Quest's printed URINALYSIS REFLEX heading as an alias
of the complete urinalysis record; 2.82.5 adds each order-set test's printed
order number as a heading surface, Quest:5616 and a bare 5616 when no
other catalog test carries it, and records members_pending on the two
memberless practice panels; members and defaults are unchanged).
assets/umls/2026AA/umls.sqlite3 # protected rebuild/audit only
assets/umls/2026AA/manifest.json # protected rebuild/audit only
The LOINC CSV files under data/ are immutable build inputs. Runtime mapping
uses the compiled schema-6 SQLite catalog, not a CSV scan. The catalog imports
release-derived COMMON_TEST_RANK and ORDER_OBS; the former is a bounded
ranking prior and the latter prevents order-only terms from representing a
result row. The FAISS file is derived
from the same LOINC release and SapBERT model recorded in terms.json. The
UMLS database is licensed material and must be provisioned through a protected
artifact store or deployment secret volume, never committed to Git.
UniversalSemanticIndex is built in memory once from this pinned SQLite
catalog when the worker constructs a Mapper; it has no separate artifact and
does not duplicate LOINC authority. It uses Components, Parts, release aliases,
consumer/related names, and axis facts to reach ordinary source-neutral
candidates before UMLS/FAISS recall.
Workers prefer serving/umls_loinc.sqlite3 when present. It is the compact
runtime artifact; retain raw UMLS only in protected rebuild/audit storage.
Code-vector files avoid repeated candidate embedding during ranking. Class
shards are optional latency artifacts, never another terminology authority.
config/axis_context_aliases.json is required runtime configuration, not
documentation-only data. Its entries provide reviewed axis surfaces and its
method_families section defines editable method equivalences used by context
extraction and safety validation. When changing a family, increment the file's
version, run the regression suite, rebuild the catalog if the release artifact
embeds the axis tables, and restart workers with the matching config bundle.
build_mapper(..., context_aliases_path=...) and the CLI
--context-aliases option let the main project provide an explicit path when
the config is stored outside the package source tree. Do not add a code target
or a force-exact rule to this file.
config/axis_safety_policy.json is release-controlled unit and reusable axis
safety configuration. config/unit_surface_rules.json holds grammar-level OCR
unit repairs. A valid dimension/property conflict remains a hard failure; an
unfamiliar raw unit is retained in unit_assessment and is reviewable rather
than silently reinterpreted.
config/panel_context_index.2.82.json (version 2.82.11) is generated from
the clinician default CSV, the local pinned LOINC PanelsAndForms.csv, the
aliases file and, since 2026-09-22, the practice's Elation order-set catalog
and the active registry (see Panel context and
CLI build-panel-index). Rebuild it whenever any input or the LOINC
release changes; its sources record every input's hash. It is a context
artifact, not a result-code registry and not a web-scraped panel list.
config/laboratory_vendors.json (version 2026.09.22.1) maps the practice's
Elation vendor ids to the names those vendors print on a fax, so a
source-scoped registry entry published as Quest is admitted for a report
that says Quest Diagnostics (see Clinician learning
"Laboratory names"). The Mapper loads it from its config directory and the
runtime manifest lists it; editing it bumps the version.
Closed-set Gemini Panel Classification
The generated panel artifact exposes a closed enum at
panel_context_enum.values[].value. Give Gemini only those canonical values
plus UNKNOWN; do not ask it to invent a panel name or parent LOINC code.
Always preserve the original report heading separately in the report/OCR
artifact.
Use this precedence in the main project:
- Match the raw heading locally against
panels[].aliases. An exact match is sent aspanel_context_source="official_panel_heading"with confidence0.98; the mapper may use documented component membership and panel defaults after normal safety validation. - If the deterministic match fails, Gemini may choose one enum value or
UNKNOWN. For an enum selection, send the canonical value withpanel_context_source="llm_enum_hint"and cap confidence at0.70. This is retrieval/ranking evidence only; it cannot activate a panel default specimen or setpanel_loinc_hint. - For
UNKNOWN, send the raw heading unchanged withpanel_context_source="unrecognized_panel_heading"and confidence0. Do not discard it and do not infer a parent code.
Example Gemini response contract:
{
"raw_panel_heading": "COMPREHENSIVE METABOLIC PANEL-CMP",
"panel_context_enum": "Comprehensive Metabolic Panel (CMP)",
"confidence": 0.94
}
The main project still performs the deterministic alias match first. It must
not promote Gemini's 0.94 to 0.98 or let Gemini write panel_loinc_hint.
Optional Gemini Unit Verification
Run assess_unit(raw_unit) first. The normal parser already handles bounded
OCR forms such as ulU/mL, u1U/mL, mgdL/, K/mcL, x10E3/uL,
x10E6/uL, and pcg, every UCUM construct the pinned release uses
(mg/g{creat}, mg/(24.h), /[HPF], [arb'U]/mL, /uL), unit names and
scale words (Thousand/uL, units/mL), report notes (mg/dL (calc),
% by wt; a print that is only a note, (calc) beside a ratio, is an
absent unit carrying the note, and a bare method word of the reviewed
vocabulary in the unit column, calc, is context, not a unit, 2026-09-25)
and the release's display forms
(arb U/mL, mg/g creatinine).
Gemini is optional and advisory only when the assessment is unrecognized
or a human needs an OCR explanation. Its response must be stored separately:
{
"raw_unit": "ulU/mL",
"verification": "likely_ocr_variant",
"suggested_surface": "uIU/mL",
"confidence": 0.93,
"reason": "capital I was likely read as lowercase l"
}
Never overwrite LabObservation.unit with the suggestion or canonical unit.
Show it in review, and apply a clinician-approved correction only as an
audited revised observation. Gemini timeout, error, or unknown leaves the
raw unit unchanged and lets normal mapper diagnostics request review.
Do not ask a clinician to invent a unit for a row that is intentionally
unitless. Before rendering a unit-correction field as required, inspect the
selected LOINC term through the review validation API. Terms whose official
example is {ratio} or {titer} can be approved with an empty report unit;
terms with a dimensional example such as mg/g, mg/dL, or mmol/L cannot.
Preflight
Run these checks in the image build or release pipeline, not for every row:
$env:PYTHONPATH = "$PWD\src"
python -m loinc_mapper validate-assets `
--umls-path assets/umls/2026AA
python -m loinc_mapper build-panel-index `
--panel-defaults config/lab_panel_default_specimens.csv `
--panels-and-forms data/Loinc_2.82/Loinc_2.82/AccessoryFiles/PanelsAndForms/PanelsAndForms.csv `
--aliases config/panel_context_aliases.2.82.json `
--output config/panel_context_index.2.82.json `
--release 2.82
python -m unittest discover -s tests -v
python -m mkdocs build --strict
The UMLS checksum check reads the entire raw SQLite file and can take minutes for the approximately 14 GB artifact. That is expected. Run it in the image build/release validation job, not worker startup. A worker validates the pinned serving manifest and file size instead of hashing raw UMLS data.
Perform one real smoke mapping and inspect its provenance:
python -m loinc_mapper map `
--name "Cholesterol, LDL, Measured" `
--value "133" `
--unit "mg/dL" `
--context-aliases config/axis_context_aliases.json `
--umls-path assets/umls/2026AA `
--catalog assets/loinc/2.82/catalog_v5.sqlite3 `
--vector-index assets/loinc/2.82/terms.faiss `
--vector-metadata assets/loinc/2.82/terms.json `
--scispacy-model en_core_sci_md
The smoke result must show semantic_index.backend=faiss, a nonzero
semantic_retrieval_candidate_count, axis_vocabulary.backend=sqlite, and
umls_retrieval_backend=serving_exact+fts5+trigram. It must not report a
NumPy vector backend or exact_only_bridge.
Run the de-identified runtime canaries against the complete production bundle before promoting an image. They catch a stale/missing registry or configuration mount for the OCR-unit, enzyme, and confirmed-panel-component regressions:
python -m loinc_mapper runtime-canaries `
--registry config/mapping_registry.order_sets.20260827.json `
--catalog assets/loinc/2.82/catalog_v5.sqlite3 `
--panel-index config/panel_context_index.2.82.json `
--safety-policy config/axis_safety_policy.json `
--umls-path assets/umls/2026AA `
--vector-index assets/loinc/2.82/terms.faiss `
--vector-metadata assets/loinc/2.82/terms.json
Persist MappingResult.diagnostics and
MappingResult.provenance.runtime_manifest with the main-project case. The
first identifies the failed stage, reason code, candidate, and recommended
next step; the second proves which registry and configuration files the worker
actually loaded.
The round-J canaries (2026-09-25) guard the 09-25 wrong codes and abstains
so a stale sync is caught at startup: direct_bilirubin (1968-7, the
released method clue and coverage), ehrlichia_igg_titer (9783-2, the
value's shape), interpretation_pointer_is_no_result (a See note row must
show value_not_a_result and no code: the one non-filing canary,
expected_loinc: null with expected_outcome), egfr_ckd_epi_no_year
(98979-8), urine_acr_creat_denominator (9318-7), absolute_basos_auto
(704-7), lymph_percent_auto_scale (736-9 with scale="%"),
testosterone_bioavailable (2990-0) and prothrombin_time (5902-2; the last
three need the approvals round 20260925 or later). Each canary case reports
primary_outcome beside status and loinc_code.
The round-K canaries (2026-09-29) fail on a package older than the review
redesign, so a stale sync stops at startup: question_on_a_result (every
result carries stages.question), titer_files_by_policy (a titer that
prints neither method nor specimen files the method-less serum term and
records policy_defaults method:method_less, system:ser/plas) and
material_tie_names_its_axis (a test family with no clear material asks,
with stages.choice). A canary may expect evidence beside its code
(RuntimeCanary.expected_stages, expected_policy_defaults); each case
reports policy_defaults and missing_evidence. The last two read the
practice's approvals: they need the 20260925 registry or a descendant that
approves neither name.
urine_sediment_count_under_urinalysis (0.3.1, 2026-10-02): RBC printed
NONE SEEN per /HPF under the trusted URINALYSIS, COMPLETE heading must
file the urine sediment count 13945-1. A package older than 0.3.1 abstains on
the practice's blood-count approval of RBC (789-8), so a stale sync fails
it at startup. It reads that approval: it needs the 20260925 registry or a
descendant that keeps it. urine_sediment_count_with_printed_urine is the
production Quest form of the same row (WBC, the section heading also prints
Urine): it must file 5821-4, which needs the trailing-word rule as well (the
count had tied with the clumps term). urinalysis_heading_after_order_number
is CPL's form: RED BLOOD CELLS under 1501 URINALYSIS W/REFLEX MICRO, which
needs panel index 2.82.9 (the heading's words, and the rule that a leading
order number is not part of a heading's name). The report has 25 cases.
expanded_abbreviation_keeps_its_approval and urine_printed_under_the_cbc_heading
(0.3.3, 2026-10-06): EGFR expanded by the model to estimated glomerular
filtration rate must still file the practice's universal eGFR approval
98979-8 (the registry is asked for the printed surface as well), and RBC
printed NONE SEEN per /HPF with Urine on the row under the trusted CBC
With Differential heading must set the blood-count approval aside and file
the sediment count 13945-1 (a printed specimen outranks the member default
for an approval the unit already refused, and the per-field unit names the
microscopy method, 0.3.4; 0.3.3 reviewed on the computer-assisted twin,
ambiguous_safe_signatures, and 0.3.2 held the approval as blocked,
approved_candidate_rejected). Both read the 20260925 registry's approvals.
nmr_total_hdl_particles, ldl_c_printed_calculated,
ldl_c_method_less_keeps_its_approval and vitamin_d_total_default (0.3.5,
2026-10-08, Release A): HDL-P (total) in umol/L under LabCorp's printed
heading NMR LipoProfile ® test must file the practice's 49748-7 (panel index
2.82.10 names the record by that heading, the approvals round 20261008 holds
the HDL-P surfaces, and the validator reads a printed fraction acronym as the
name of the dotted sub-fraction a total row states); LDL-C with the printed
method calculated must file 13457-7 (the second question of the name) while
the bare LDL-C keeps the method-less approval 2089-1; and Quest's
VITAMIN D, 25 OH must file the D2+D3 total 62292-8 (the round replaced the
D3-fraction approvals of the total prints). All four need the 20261008
registry (registry-20261008T172328-65b01b4d64f1) or a descendant, the first
one index 2.82.10 as well: a package or registry older than Release A fails
them at startup. Release B of the same package adds
synonym_slash_files_the_approval (SGPT/ALT in U/L under the CMP heading
must file the practice's 1742-6: the slash between two names of one enzyme
enumerates), unit_less_row_reads_the_practices_property (Creatinine,
Urine printing a number and no unit must file 2161-8, the Property the
practice approved the analyte in, flagged unit_missing) and
urine_microscopy_word_reads_presence (RBC printing FEW with Urine
under the urinalysis heading must file the presence term 32776-7). The
sandbox round (0.3.6, 2026-10-09) adds omegacheck_fraction_is_whole_blood
(EPA in % by wt under Quest's OMEGACHECK(R) heading must file the Blood
fraction 90912-7: index 2.82.11 puts the record on LOINC's Blood panel and
the approvals round 20261009 makes the practice's Blood fractions universal),
nucleated_rbc_per_100_leukocytes_is_the_ratio (NUCLEATED RBCS printed
/100 WBC'S must file the ratio 58413-6, never the count per volume) and
venous_lead_in_mcg_per_dl (LEAD (VENOUS) in mcg/dL must file 77307-7).
They need the 20261009 registry
(registry-20261009T163241-e68ab7715a95) or a descendant and index
2.82.11; the unit-grammar list gains /100 WBC'S -> /100.{WBC}. The report
has 37 cases.
runtime_manifest.package (additive, 2026-09-25) proves which mapper code
produced a case: name, version (loinc_mapper.__version__, 0.2.0 for
round J, 0.3.0 for round K) and build, the stamp the hot-sync script writes to
src/loinc_mapper/_build.py (git_sha, tree_hash, synced_at, by; the
file is generated, gitignored, and build is null when the tree carries
none). The package is loaded from a bucket at runtime, so an image tag proves
nothing about it; a manifest without package predates round J.
For the universal layer, also inspect
result.provenance["stages"]["universal_axis_facts"] and
result.provenance["stages"]["axis_signature_groups"]. A normal urine,
CSF, 24-hour, calculated, or direct-assay row should show the report facts
that removed its incompatible LOINC siblings.
result.provenance["stages"]["registry_evidence"] shows separately whether a
universal template or source-specific record participated. routing_behavior
shows that class metadata was advisory and laboratory-wide retrieval remained
available; it is useful when an OCR section label is wrong.
Registry Approval Replay
Clinician approvals are published as authority=clinician_override, but this
authority clears only semantic confidence and margin after normal UCUM,
active-code, six-axis, context, and exact-alias checks pass. It is not a force
map. The main project uses the fast publication policy for clinician-approved
universal templates:
policy = PublicationPolicy.CLINICIAN_FAST_REPLAY
candidate = compile_registry_snapshot(
active_registry,
approvals,
catalog,
publication_policy=policy,
)
if candidate.requires_publication:
# Factory builds with PipelineSettings.clinician_review_replay().
replay = replay_registry_snapshot(
candidate,
approvals,
mapper_factory,
publication_policy=policy,
progress=record_progress,
)
active = finalize_registry_snapshot(
candidate,
replay,
publication_policy=policy,
progress=record_progress,
)
This performs one deduplicated map_many() batch over the original reviewed
rows. The review mapper loads no UMLS, ScispaCy, FAISS, SapBERT, or learned
ranker assets. It still performs deterministic context, exact alias, active
term, UCUM, and six-axis checks. A rejected unit/specimen/method/time/scale
cannot publish just because a clinician selected a code.
PublicationPolicy.STRICT_REPLAY remains the scheduled release-audit mode. It
performs full original/unseen and fast equivalence replays, but it must not be
used synchronously from a clinician browser action.
Publication generates bounded verified_aliases from the approved name for
lossless/compact and reviewed OCR variants. These variants remain linked to
the approval and are safe only after replay and normal validation. A contained
or arbitrary fuzzy registry hit does not inherit clinician authority.
If the active registry already contains legacy source evidence for the same alias and target, the publisher should not delete it. The compiled snapshot adds the approved universal template, and runtime candidate merging prefers the clinician-approved universal evidence for that code. If the legacy entry has the same scope as the approval, compilation upgrades that entry in place.
Workers must load only a finalized registry_status=active snapshot. Candidate
snapshots and browser-submitted registry JSON must be rejected. The active
pointer should be advanced with a Cloud Storage generation precondition so two
publishers cannot overwrite one another. The loaded MappingRegistry exposes
registry_status (the snapshot's own value) and sha256 (its content hash,
an alias of source_sha256) so a health endpoint can show which snapshot a
worker actually loaded; before 2026-09-17 those attributes did not exist and
the main project logged registry_status=None sha256=None.
Asynchronous Publisher Job
The FastAPI request persists the append-only decision and starts a private registry-publisher Cloud Run Job. It returns the review status immediately; the browser polls the main project rather than waiting on mapper work. The job:
- Loads the active immutable snapshot and only pending/superseding decisions.
- Emits
validating,batch_queued,batch_completed, andfinalization_readyprogress fromPublicationProgressEvent. - Skips replay and registry writes for
candidate.requires_publication=false, recordingno_changefor those audit decisions. - Writes the finalized immutable snapshot, advances the active pointer with a
generation precondition, then records
published. - Records
failed_safetywith bounded de-identified diagnostics when deterministic validation or the exact replay fails.
Keep strict model replays in a separate scheduled Cloud Run audit job. Those audits may create clinician-visible follow-up work but never silently roll back an already authorized clinician approval.
Deferred Order Context
Pending Elation orders are intentionally deferred to the next version. After
demographic matching, the main project may inspect the patient chart and pass
an advisory OrderContext with an order-set/test identifier, local lab test
identifier, CPT, diagnosis, and ordering context. It must not be a hard global
LOINC filter. If it is absent, stale, ambiguous, or unsafe, the mapper uses
normal laboratory-wide retrieval.
The current practice CPT CSV is advisory inventory only, not an active practice-specific subset, so it must not filter the LOINC catalog. Pasted order sets follow the same rule until the future order-context integration is clinically mapped and replay-tested.
SapBERT is loaded from the pinned local model snapshot only; production workers
do not download it from Hugging Face at request time. A Windows error such as
The paging file is too small for this operation means the host virtual-memory
configuration cannot reserve the model weights. Increase the paging file or
run the worker on a host/container with sufficient memory; do not replace
SapBERT with a smaller unvalidated model as a workaround.
Python Integration
Instantiate the mapper once during worker startup. Do not construct ScispaCy, UMLS connections, SapBERT, or FAISS inside a request loop.
from pathlib import Path
from loinc_mapper import LabObservation, build_mapper
from loinc_mapper.elation import ElationAdapter
ROOT = Path("/opt/noise-to-loinc")
VECTOR_SHARDS = ROOT / "assets/loinc/2.82/shards/vector_shards.json"
mapper = build_mapper(
core_path=ROOT / "data/Loinc_2.82/Loinc_2.82/LoincTableCore/LoincTableCore.csv",
rich_path=ROOT / "data/Loinc_2.82/Loinc_2.82/LoincTable/Loinc.csv",
common_names_path=ROOT / "data/loinc_common.csv",
registry_path=ROOT / "config/mapping_registry.json",
map_to_path=ROOT / "data/Loinc_2.82/Loinc_2.82/LoincTable/MapTo.csv",
catalog_path=ROOT / "assets/loinc/2.82/catalog_v5.sqlite3",
umls_path=ROOT / "assets/umls/2026AA",
scispacy_model="en_core_sci_md",
sapbert_model="cambridgeltl/SapBERT-from-PubMedBERT-fulltext",
vector_index_path=ROOT / "assets/loinc/2.82/terms.faiss",
vector_metadata_path=ROOT / "assets/loinc/2.82/terms.json",
vector_shards_path=VECTOR_SHARDS if VECTOR_SHARDS.exists() else None,
panel_index_path=ROOT / "config/panel_context_index.2.82.json",
safety_policy_path=ROOT / "config/axis_safety_policy.json",
)
observation = LabObservation(
raw_name=ocr_row.name,
value=ocr_row.value,
unit=ocr_row.unit,
specimen=ocr_row.specimen,
collection_specimen=ocr_row.collection_specimen,
method=ocr_row.method,
time_context=ocr_row.time_context,
scale=ocr_row.scale,
reference_range=ocr_row.reference_range,
source_laboratory=ocr_row.source_laboratory,
source_test_id=ocr_row.source_test_id,
observation_domain=ocr_row.observation_domain or "LAB_RESULT",
loinc_class_hint=ocr_row.loinc_class_hint,
panel_context=ocr_row.panel_context,
panel_context_source=ocr_row.panel_context_source,
panel_context_confidence=ocr_row.panel_context_confidence,
panel_loinc_hint=ocr_row.panel_loinc_hint,
report_section_id=ocr_row.report_section_id,
class_source=ocr_row.class_source,
class_confidence=ocr_row.class_confidence,
report_context=ocr_row.report_context,
)
result = mapper.map(observation)
if result.status == "mapped":
payload = ElationAdapter().to_payload(observation, result)
# Send payload to Elation only after recording result.provenance.
else:
# Store result.to_dict() in the review queue. Never guess a LOINC code.
payload = None
When diagnosing an unexpected abstention, verify the worker loaded this workspace's package and registry rather than an older installed copy:
import inspect
import loinc_mapper
print("loinc_mapper:", inspect.getfile(loinc_mapper))
print("registry:", mapper.registry.version)
probe = mapper.map(LabObservation(
raw_name="SEX HORMONE BINDING GLOBULIN",
value="162.00",
unit="nmol/L",
specimen="Blood",
observation_domain="LAB_RESULT",
panel_context="TESTOSTERONE, TOTAL AND FREE",
))
print(probe.status, probe.loinc_code)
print(probe.provenance.get("mapping_registry_version"))
With the reviewed endocrine aliases, the probe should return mapped and
13967-5. A worker that reports an older registry version, imports the main
project's legacy faxautomation lookup, or returns abstain has not loaded
the current mapper artifact. Restart or replace that worker after updating the
package and its explicit immutable active-registry snapshot.
For a fax with several rows, use mapper.map_many(observations). This reuses
the release-derived universal index, UMLS lookups, FAISS query embeddings,
SQLite access, and process memory while keeping every row's safety decision
independent.
Close the mapper when a short-lived job exits so Windows and container runtime file handles are released cleanly:
try:
results = mapper.map_many(observations)
finally:
mapper.close()
OCR reconstruction boundary
Faxautomation must repair report tables before constructing LabObservation.
Preserve original OCR cells privately, call Gemini once only for a damaged
table region, and send a separate mapping_name with bounded repair evidence.
Never add transient surfaces such as Glucose 02 to the registry. Send
input_quality="input_structure_invalid" for merged, shifted, duplicate, or
orphan rows; the mapper will abstain before semantic retrieval. The complete
payload, Gemini JSON contract, panel rules, and force-map queue behavior are
defined in OCR recovery contract.
If the main project already owns OCR table extraction, use its parser and pass
rows through the adapter above. Use map-report only when the input is a text
report and the package's replaceable report extractor is appropriate.
Result Contract
Persist the complete MappingResult.to_dict() for audit and review. At a
minimum, retain:
status,loinc_code,confidence, andmargin.- normalized input and normalization operations.
- candidate codes and retrieval sources.
- six-axis and unit evidence, including hard rejections.
- UMLS, FAISS, model, catalog, registry, and release provenance.
- abstention or rejection reasons.
primary_outcome.reason_codemay beunnamed_top_candidate(additive, 2026-09-24): the top-ranked term shares no word with the row while another candidate, named incandidate_code, does; the row abstains for review.stages.registry_evidenceitems carryregistry_match(exact,joined,contained,contained_unexplained, ...) andexact_for_row(additive, 2026-09-24): a consumer tells an approval of the row's surface from a retrieval hit without re-deriving it.primary_outcome.reason_codemay bevalue_not_a_result(additive, 2026-09-25): the value cell is a pointer (See note), a status (TNP) or a declared non-result (result_kind); no candidates, nothing to file, keep the cell for the note.stages.value_shapenames the shape; the cell text is not echoed.stages.universal_axis_facts.valuecarries the shape of every value (number,titer,answer, ...), andinapplicable_approved_candidatesnow also lists an approval rejected on the value's shape (shape_conflictin its validation evidence).primary_outcome.reason_codemay bemethod_variant_ambiguity(additive, 2026-09-25): the leading candidate, named incandidate_code, differs from the alternatives of its meaning group only in a method the row does not print (stages.meaning_groups[].unresolved: "method_variant"); the row's method column or the practice's approval of one variant decides.stages.policy_defaults(additive, 2026-09-29): on a mapped row, the choices the practice's policy made where the report printed nothing:{axis: "method", choice: "method_less", code, completed?}and{axis: "system", choice: <material>, code, class}. Absent when the report printed every fact. A filing note can say "no method printed: the general code" from it; the material default also keeps the review flagspecimen_unprinted.stages.choice(additive, 2026-09-29): on a row that reviews on a tie (primary_outcome.reason_codeambiguous_safe_signaturesormethod_variant_ambiguity), what a card prints:{options: [codes], separating_axes: [{axis, values: {code: words}}], printed: {kind, unit_family, reference_range_shape}, policy_choice}.axisis LOINC's own axis name (Component,Property,Time,System,Scale,Method), a value an option does not state readsnot stated,printednames the shape of what the report printed and never the value, andpolicy_choiceis the code the policy would preselect (the leader). The options are the leader, the unresolved siblings of its group and the runner-up inside the margin, five at most.diagnostics[].reason_codemay beexpansion_rejected(additive, 2026-09-25, non-blocking, stageinput_repair): the application'sllm_name_expandedexpansion was not spelled by the print's letters (grounding.abbreviation_covers), the print was mapped, andstages.input_repair.repair_operationsends withexpansion_rejected;batch_counters.expansion_accepted/expansion_rejectedcount them.stages.panel_context.panel_name_candidates(additive, 2026-09-25): the index records whose members the row's own name names, as[{panel_key, default_specimen}]; evidence for the consumer's membership checks and for the row's default material, never a hard panel.stages.released_clues(additive, 2026-09-25): method or time clues read out of the row name that LOINC itself uses to name a retrieved candidate's analyte (directinDirect Bilirubin), voided for the row:[{axis, value, span, named_by}], empty when none.stages.coverage_dominance(additive, 2026-09-25):null, or{leader, displaced, coverage}when a group winner whose names explain strictly more of the row's words than the ranked leader's took the lead with its own score;displacedlists the winners it strictly covers, moved to the alternatives.review_flags(additive, 2026-09-24): why a mapped row still wants a reviewer's eye, becausereasonsis emptied on a mapped row. Todayunit_incomplete,unit_unreadableandunit_prefix_unreadable: an exact clinician approval or confirmed panel member proceeded on a unit that states only a magnitude, could not be read, or lost its SI prefix; andunit_missingfor any candidate that went without the unit its code would need (units never block since 2026-09-24); andspecimen_unconfirmed: an approval of the row's own surface mapped although a specimen not printed on the row (areport-level phrase) contradicts its System; andspecimen_unprinted(2026-09-25): the row printed no material and the practice never approved the analyte, so the term in the default material of its test family led its System siblings. File the row withneeds_loinc_review=True; the code is right, the printed unit or the material needs a human.
Only status == "mapped" may create an Elation test.loinc value by default.
Treat abstain, no_candidate, and invalid_candidate as review work, not as
failed API calls to retry blindly. force_mapped is a separate audited
operational exception: it remains blocked from Elation unless trusted server
deployment configuration explicitly sets IS_FORCE_MAPPED_ELATION_ON=true.
Never let a browser, OCR row, or clinician form set that deployment switch.
For the certified OCR-recovery snapshot, leave
IS_FORCE_MAPPED_ELATION_ON=false. Send every force_mapped output to the
authorized final-approval queue together with its raw OCR cells, repair trace,
result style, matched force alias/target, and mapper provenance. The three
Labcorp-specific force entries require source_laboratory="Labcorp"; do not
invent a vendor name when it was not extracted.
Send raw unit text exactly as printed. There is intentionally no finite UI unit
enum because UCUM expressions are composable. The result returns a structured
unit_assessment with valid, recovered, missing, or unrecognized
state; use that state to guide review without overwriting the raw report text.
Its keys are raw, surface, canonical, dimension, state, valid,
reason, and, since 2026-09-17, recovered_surface (the accepted surface of
a recovered repair) and annotations (report notes such as calc and UCUM
annotations such as creat; a list, never part of the dimension). The
reason literal for a repair, recovered OCR unit surface: X -> Y, is
unchanged.
Never send unit_assessment.canonical back as LabObservation.unit: it is
validation evidence, not OCR text. The mapper accepts its own {arb}/L
canonical arbitrary-unit form for backward compatibility with earlier
integrations, but the main project must preserve and display the fax surface,
for example U/L.
If the main project uses the DSPy review adapter in temp_file/judge.py, pass
the complete MappingResult.to_dict() and the original row context. The
judge may explain or suggest a safety-approved candidate for a reviewer, but
the main project must still use the mapper status/code for filing. Never use
loinc_judge_code as an Elation code when the mapper status is not mapped.
Persist the judge decision and reason beside the mapper provenance for audit.
Configure LOINC_CATALOG_PATH in the main project to the same pinned
assets/loinc/2.82/catalog_v5.sqlite3. The judge independently checks active
status, class, UCUM dimension, specimen, and all six axes. It fails closed if
the catalog is missing. Do not let a judge code become an Elation code.
The OCR adapter must preserve duplicate analyte rows. A dictionary keyed only by name will lose a second Calcium, bilirubin, or hormone row. Use a list of row objects with page/report identity and pass one LabObservation per row. Pass collection-only text via collection_specimen; do not claim that a venous collection proves the analytical material is serum/plasma or whole blood.
Clinician Review, Registry Learning, and Publishing
The production review path is described in detail in Clinician-governed learning. The main project must persist review cases, case revisions, mapper provenance, clinician decisions, reviewer identity, status history, report references, and registry-version references in a durable ledger. This package deliberately has no frontend, HTTP routes, Firestore client, Cloud Storage client, or patient-data database.
Staff can draft or correct facts, while an authorized clinician approves a
reusable mapping with rationale. The UI must show the original OCR row,
report/page link, mapper reason, safe candidates, active LOINC details, six
axes, and example units. Do not expose raw registry JSON,
fast_path_enabled, or ordinary force-exact controls to clinicians. A
restricted server-side operational workflow can create an audited
approve_force_exact decision only after policy approval; it is not a clinical
confidence override and never silently enables Elation filing.
The application boundary is:
review_case = build_review_case(observation, result, report_reference)
validation = validate_review_decision(review_case, decision, catalog, active_registry)
policy = PublicationPolicy.CLINICIAN_FAST_REPLAY
candidate = compile_registry_snapshot(active_registry, approvals, catalog, publication_policy=policy)
if candidate.requires_publication:
replay = replay_registry_snapshot(candidate, approvals, mapper_factory, publication_policy=policy)
active_snapshot = finalize_registry_snapshot(candidate, replay, publication_policy=policy)
correct_and_recheck creates an audited correction and remaps the row. It does
not create a registry mapping. not_standardized and not_lab_result are
manual outcomes that never emit test.loinc. The clinician-fast policy runs
one original-row map_many() replay using
PipelineSettings.clinician_review_replay() and immediately enables the exact
universal fast path after it passes. Source-specific entries remain
full-pipeline only. Clinician authority clears confidence and margin only after
exact-alias, active-code, UCUM, and six-axis validation; it is not a force-map.
Use PublicationPolicy.STRICT_REPLAY only in an asynchronous audit/release
job. It retains original/unseen full-pipeline replay and fast equivalence
checks without blocking clinician publication.
The candidate snapshot has registry_status=candidate, so mapper workers
reject it. Only the finalized replay-passing payload is active and can be
published as a new immutable registry object.
Legacy JSONL Diagnostics
The following compatibility commands are for local diagnostics or migration of an old JSONL queue. New clinician approvals should use the review-case contract and CSV workflow instead. The compatibility importer never enables the registry fast path.
python -m loinc_mapper review-queue `
--input observations.csv `
--output evaluation/review_queue.jsonl `
--umls-path assets/umls/2026AA
python -m loinc_mapper import-reviews `
--input evaluation/review_queue.approved.jsonl `
--output config/mapping_registry.reviewed.json
Review data must identify the reviewer, decision, selected active LOINC code, and rationale for unit, specimen, method, and other axis choices. Registry updates are vocabulary changes, not automatic model retraining. Promote them through normal review and deployment controls.
When an expert approves a review row, set review.mapping_scope explicitly:
universal_templateonly when the display means the same clinical measurement across laboratories and its unit/axis constraints are known.source_evidencefor a local identifier, vendor assay, or report-specific display that needs the named source laboratory.
For a universal approval sourced from one laboratory, the importer records
that laboratory as supporting_source_laboratory in provenance rather than
making it an execution-time requirement. Audit and export the two layers
without changing the live registry:
python -m loinc_mapper audit-registry-scopes `
--registry config/mapping_registry.json `
--output evaluation/registry_scope_audit.json
python -m loinc_mapper export-layered-registry `
--registry config/mapping_registry.json `
--output config/mapping_registry.layered.json
Cluster recurring abstentions before requesting clinical review:
python -m loinc_mapper cluster-review-queue `
--input evaluation/review_queue.jsonl `
--output evaluation/review_clusters.json
Runtime manifest: the question key (2026-09-29)
runtime_manifest.question_key names the version of the review question's fold
(version, the package's question.KEY_VERSION) and the inputs it depends on (the
package's OCR fold and the configurations fingerprinted under configs). The main
project stores the version on every review case; a case with another version is
treated like an old-format file.
Runtime manifest: the registry in force (2026-09-29)
runtime_manifest.registry reports the registry the mapper holds:
| Key | Content |
|---|---|
version |
the registry's version; for a pending view <active version>+pending.<applied count>.<first eight of the fingerprint> |
status |
the registry's own status: active for a published snapshot, pending_view when approved, unpublished decisions apply (it was a fixed active before) |
base_version |
the published registry the view stands on; equal to version for a published snapshot |
pending |
{count, fingerprint}: the pending entries in force and pending_fingerprint(base_version, decision_ids); 0 and null for a published snapshot |
logical_path, sha256 |
the published snapshot's file and hash; a view carries its base's |
A health endpoint proves which decisions apply by comparing pending.fingerprint with
the fingerprint of the decisions it holds as approved. A pending view is installed as
an object (Mapper.replace_registry(view)); it is never written to the asset bucket,
MappingRegistry.from_json refuses its payload and a publish never takes it as its
parent. Round K needs no new registry, catalog or panel index: the package and
config/axis_safety_policy.json (2026.10.05.1) change together; Release B
(0.3.5, 2026-10-08) bumps the policy to 2026.10.08.1 (property_default,
scale_default), and config/ and src/loinc_mapper sync together as always.
Operational Safeguards
- Load one immutable artifact bundle per worker and expose its versions in health metadata.
- Do not send an unmapped or abstained result to Elation with a guessed code.
- Keep UMLS files and patient data out of logs, Git, container layers, and error reports.
- Set a resource limit for concurrent mappings; SapBERT is CPU/GPU and memory intensive.
- Use batch mapping for reports and avoid parallel workers competing for the same large SQLite file during artifact builds.
- Alert on increased abstention rate, dangerous unit/property rejections, missing FAISS evidence, or a UMLS backend downgrade.
- Retain enough provenance to reproduce a decision after a release update.
Batch Worker and Cloud Run Job
Use MappingBatch instead of starting a mapper for every fax. Group incoming
reports for up to 60 seconds or 20 reports / 500 observations, then submit one
Cloud Run Job task. The task creates one mapper process, maps the whole batch,
persists the result, and exits.
{
"schema_version": "1",
"batch_id": "fax-batch-001",
"registry_snapshot": {
"registry_version": "registry-20260807T0100-abc123",
"uri": "gs://private-bucket/registries/mapping_registry.registry-20260807T0100-abc123.json",
"generation": "1722980000000000",
"sha256": "..."
},
"reports": [
{
"report_id": "fax-123",
"observations": [{"raw_name": "Hb A1c", "value": "6.5", "unit": "%"}]
}
]
}
python -m loinc_mapper run-worker `
--input batch.json `
--output batch-result.json `
--umls-path assets/umls/2026AA
Start Cloud Run Jobs with one task at a time, 4 vCPU, 8 GiB, and
scale-to-zero. Set MAPPER_CPU_THREADS=4 so Torch, FAISS, and BLAS do not
oversubscribe the worker. Increase memory only from measured peak RSS. The
main fax service owns callback delivery, retries, idempotency, and durable
mapping/review status; persist the mapper result before Elation submission.
At job start, the launcher downloads the named active registry snapshot to
ephemeral local storage, verifies its checksum, constructs Mapper from that
file, and passes the same reference in MappingBatch. BatchWorker refuses
to run if the loaded registry version differs from the batch reference. This
lets already-running jobs finish against their explicit old snapshot while new
jobs use a newly approved one.
Keep PDFs, OCR input/output, mapper results, and registry snapshots in private Cloud Storage. Keep raw licensed UMLS build data in a protected build/audit location; package only the compact serving artifact and pinned model/catalog bundle in the private worker image. Do not bake raw licensed data or mutable registry content into a public image.
Use separate least-privilege service accounts:
- Main API: read/write review ledger and report storage; trigger Jobs.
- Mapper Job: read its batch and finalized registry snapshot; write mapping results; call an authenticated internal completion route.
- Registry publisher: read approved review events; write versioned snapshots and atomically update the active pointer.
Use immutable Cloud Storage object names and a generation precondition when updating the active pointer. See Cloud Storage request preconditions and Cloud Run Jobs. If using Firestore for the review ledger, deployment still requires the organization's executed BAA, IAM controls, audit logging, retention policy, and security review; see Google Cloud HIPAA covered products.
Serving-Artifact Build
Build the UMLS serving index in a separate protected build job. It validates the raw release checksum once and can take substantial temporary disk space; it does not modify the raw UMLS index. On Windows, launch it in the background and follow the progress log on standard error:
$env:PYTHONPATH = "$PWD\src"
$env:PYTHONUNBUFFERED = "1"
New-Item -ItemType Directory -Force evaluation\logs | Out-Null
$job = Start-Process `
-FilePath "$PWD\.venv\Scripts\python.exe" `
-ArgumentList @(
"-u",
"-m", "loinc_mapper",
"build-umls-serving",
"--asset-path", "assets\umls\2026AA",
"--catalog", "assets\loinc\2.82\catalog_v5.sqlite3",
"--output", "assets\umls\2026AA\serving\umls_loinc.sqlite3"
) `
-WorkingDirectory "$PWD" `
-RedirectStandardOutput "evaluation\logs\umls_serving_build.out.log" `
-RedirectStandardError "evaluation\logs\umls_serving_build.err.log" `
-WindowStyle Hidden `
-PassThru
$job.Id
Get-Content evaluation\logs\umls_serving_build.err.log -Wait
After it exits, run validate-assets once in the build/deployment pipeline and
confirm a smoke mapping reports serving_exact+fts5+trigram. Do not deploy the
new artifact until the profile matrix and safety regression checks pass.
Pipeline Profiles
PIPELINE_PROFILE=production requires all evidence stages. It permits an
early result only for a fully validated verified_active exact registry entry.
PIPELINE_PROFILE=evaluation permits controlled ablations but marks results
experimental, and ElationAdapter rejects those results. Use the profile
matrix on a frozen expert-reviewed corpus before changing a production stage or
the registry fast-path flag.
An approval response with eligible_for_publish=true means the decision passed
package validation and was stored by the main project. It does not mean the
worker is using it yet. The publisher must write the immutable candidate, run
the selected replay policy, finalize it as registry_status=active, atomically
advance the active pointer, and restart or reload workers. If a fax row has
only collection context, use the published clinician universal template; do
not copy Bld or Ser/Plas into the raw specimen field unless the report
actually states the analytical material.
Deployment Promotion
Use a new versioned artifact directory for every LOINC or UMLS release. Run tests, asset validation, smoke mappings, and the expert holdout evaluation before changing the worker's active configuration. Promote by changing the artifact pointer/configuration and restarting workers. Roll back by restoring the previous pointer; do not edit a live SQLite file or delete the previous bundle while workers may still have it open.