Performance, Profiles, and Worker Batches
The mapper is intentionally safety-first. Performance work removes duplicated computation; it never removes the clinical checks that decide whether a LOINC code can be emitted.
Production Flow
For every observation that is not a verified exact universal-template fast-path case:
normalize
-> AxisFactExtractor + UniversalSemanticIndex
-> deterministic SQLite retrieval
-> ScispaCy mention and abbreviation processing
-> local UMLS CUI retrieval
-> FAISS/SapBERT semantic retrieval
-> UCUM and six-axis validation
-> feature ranking
-> confidence and margin gate
The only permitted shortcut is an exact reviewed universal_template alias
with review_status=verified_active and fast_path_enabled=true. It still proves:
- the target code is active in the pinned LOINC release;
- declared unit, class, property, method, and specimen requirements are met;
- UCUM/unit and six-axis validation pass;
- no inferred context conflicts exist; and
- the confidence policy accepts the single reviewed target.
Fast-path provenance lists skipped semantic stages. A contained alias, legacy registry row, source-specific evidence record, unknown unit, or missing required specimen/method always runs the full pipeline.
Settings
Copy .env.example to the ignored .env file and keep production settings
enabled:
PIPELINE_PROFILE=production
IS_UMLS_ON=true
IS_FAISS_ON=true
IS_SAPBERT_ON=true
IS_VALIDATION_ON=true
IS_CONFIDENCE_ON=true
IS_REGISTRY_FAST_PATH_ON=true
MAPPER_CPU_THREADS=4
production rejects any disabled evidence stage. evaluation is the only
profile that allows controlled ablations. Any evaluation result is marked
execution.experimental=true, and ElationAdapter refuses to create a
payload from it.
Validation and confidence cannot be disabled in any profile. Do not use a profile experiment as a deployment setting.
Why It Is Faster
- The SQLite catalog loads once per worker. Class/type membership and token statistics are kept in memory for O(1) scope checks.
- The release-derived universal component/axis index is compiled once from that loaded catalog. It creates source-neutral candidates from facts such as urine, CSF, 24-hour duration, and stated methods; it does not query a per-laboratory mapping table for ordinary rows.
- Deterministic lookup uses the indexed
alias_tokenstable first. FTS5 is a bounded prefix fallback, not a broad query over every matching alias. - Character recovery uses distributed 3-of-4 trigram intersections. It still
tolerates a single damaged OCR trigram, but avoids the old broad trigram OR
scan for words such as
cholesterol. - Deterministic candidates are generated once per verbatim row name within a batch (the evidence depends on the exact string), and context extraction is cached per normalized name. Text normalization, unit parsing and the semantic query variants are memoized process-wide: they are pure functions of the pinned vocabularies, and the validator calls them about ten times per candidate.
- Alias similarity is gated by
SequenceMatcher's two upper bounds (normalization.best_alias_ratio,ratio_if_at_least): the exact ratio is computed only when it can still beat the best score or reach the threshold, and ties resolve exactly as the ungated scan did. - ScispaCy processes a report through
nlp.pipe; parser/tagger components not needed for linking are disabled. - The UMLS linker owns one read-only SQLite connection and caches mention-to-CUI results for the worker lifetime; exact lookups and the fuzzy decision stay per mention, while the CUI-to-LOINC bridge is queried once per batch of mentions.
- The result cache is keyed on the routing projection of
report_context(resolve_routingis its only mapping-path reader), so identical rows of different reports share a cached answer while a job id or filename never changes one; cached results are stored as pickled bytes and the runtime manifest is copied by JSON parse, both value-equal and isolated from caller mutation. - A linked UMLS name receives bounded exact/token local expansion rather than replaying the complete OCR deterministic pipeline once for every sibling CUI. Direct CUI-to-LOINC bridge candidates are always retained.
- SapBERT encodes unique queries in batches. FAISS searches are batched too.
- Code-level SapBERT vectors avoid re-embedding candidate descriptions during reranking.
- Optional class shards search an official LOINC class first, then use a laboratory/global fallback if a shard produces no scoped hit.
Every result records provenance.stages.timings_ms, candidate counts, profile,
artifact versions, cache hits, and fallback evidence. Inspect those fields
before changing a threshold.
timings_ms covers only the per-row stage (_map_one). The batch prologue
that map_many runs once per report is reported separately, with the same
values copied onto every result of that batch:
provenance.stages.batch_timings_ms—panel_resolve_ms,row_probes_ms(cache, force-exact and fast-path probes),deterministic_batch_ms,umls_link_batch_ms(ScispaCynlp.pipeplus per-mention UMLS lookups),faiss_prefetch_ms(query encoding plus the batched FAISS search),map_one_rows_ms,total_ms.provenance.stages.batch_counters— the delta of process-lifetime counters for that batch:result_cache_hits,force_exact_results,fast_path_results,pending_rows,structure_invalid_rows,umls_mentions_linked,umls_mention_cache_hits,umls_fuzzy_fallbacks(+umls_fuzzy_ms),umls_trigram_fallbacks(+umls_trigram_ms),semantic_cache_hits,semantic_cache_misses,catalog_fuzzy_queries(+catalog_fuzzy_ms),embedding_cache_hits,embedding_encoded(+encode_ms),faiss_queries(+faiss_search_ms).provenance.stages.batch_size.
When a report's wall clock is far above the sum of its rows' total_ms, the
difference is in batch_timings_ms; the counters say which fallback paid.
Measuring a change
Measured on 2026-09-17 with the production snapshot and force overrides
stripped: the cold pass fell from 260.1 s to 132.2 s for the 55 rows after the
results-identical round (memoized normalization and unit parsing, gated alias
similarity, batched UMLS bridge, routing-keyed result cache);
validation_ms mean 712 -> 69 (p95 1748 -> 169), umls_candidate_union_ms
mean 4320 -> 2448, deterministic_batch_ms 11.9 -> 8.4 s. The remaining
cold cost is the per-concept generate_semantic misses inside
umls_candidate_union_ms, the next target. Pass 2 of the benchmark clears the
result cache first (the routing-keyed cache would otherwise serve every row of
a repeated fixture in well under a second, which is the production case for
recurring rows but not a pipeline measurement); --keep-result-cache shows
the cache-read case instead.
tools/benchmark_map_many.py is the operator gate for performance work. It
maps the 55-row de-identified fixture (tests/fixtures/ocr_recovery_review_cases.20260910.json)
through build_mapper on the real artifact bundle, strips the force-exact
overrides by default so the certified rows exercise the full pipeline, and
prints the batch timings, counters, and per-stage p50/p95 for each pass:
$env:PYTHONPATH = "$PWD\src"
python tools\benchmark_map_many.py --umls-path assets/umls/2026AA --repeat 2 `
--write-results evaluation/runs/baseline.json
# after the change
python tools\benchmark_map_many.py --umls-path assets/umls/2026AA --compare evaluation/runs/baseline.json
--compare exits 1 when any row's mapping identity (status, code, ordered
candidates, confidence, margin, reasons, axis and unit evidence, diagnostics)
differs from the saved run. Pass two shows warm-cache behaviour. Results go
under the gitignored evaluation/runs/; the tool writes nothing under
config/. On a Windows host, a SapBERT OSError 1455 (paging file too small)
is an infrastructure limit, not a reason to substitute a smaller model.
Historical note: the single-row deterministic retrieval of
VLDL Cholesterol Cal measured 611.6 ms after the alias-index fix (down from
40-48 s). End-to-end numbers come from the benchmark above and from the
production batch_timings_ms.
The raw 2026AA fallback now performs intersection-based token lookup and
measured 5.3 s for that same UMLS name, down from 17.3 s; it is still too
slow for the report target. The compact serving artifact remains a required
deployment gate for production latency measurements.
UMLS Serving Artifact
The raw UMLS umls.sqlite3 may be large because it contains broad
Metathesaurus and older trigram structures. Build this compact runtime artifact
once after verifying raw assets:
python -m loinc_mapper build-umls-serving `
--asset-path assets/umls/2026AA `
--catalog assets/loinc/2.82/catalog_v5.sqlite3 `
--output assets/umls/2026AA/serving/umls_loinc.sqlite3
It retains UMLS aliases only when their CUI bridges to an active local
LOINC code and creates exact, FTS5, and trigram indexes. The source release is
validated by checksum during this build. At worker startup, the serving
manifest validates release, artifact shape, and file size without rehashing
the raw 14 GB file. Runtime provenance should then show
serving_exact+fts5+trigram.
Keep raw licensed data private and immutable so the serving artifact can be rebuilt and audited later.
For the background PowerShell build, follow progress on standard error:
Get-Content evaluation\logs\umls_serving_build.err.log -Wait
FAISS Derived Artifacts
An existing terms.faiss can gain code-level vectors without a new SapBERT
encoding run:
python -m loinc_mapper build-code-vectors `
--index assets/loinc/2.82/terms.faiss `
--metadata assets/loinc/2.82/terms.json
Optional laboratory/class shards are also derived from the same global index:
python -m loinc_mapper build-vector-shards `
--catalog assets/loinc/2.82/catalog_v5.sqlite3 `
--index assets/loinc/2.82/terms.faiss `
--metadata assets/loinc/2.82/terms.json `
--output assets/loinc/2.82/shards
Shards are a latency optimization only. SQLite remains the authoritative LOINC catalog, and every shard result still undergoes the same laboratory-type, UCUM, and six-axis safety checks. Class agreement is advisory evidence; a laboratory-wide fallback prevents a mixed report section from erasing a valid candidate. Benchmark class shards before shipping them because they consume additional disk space.
Batch Worker Contract
The main fax project should submit up to 20 reports or 500 observations to one
long-lived worker process. MappingBatch version 1 looks like this:
{
"schema_version": "1",
"batch_id": "fax-batch-2026-07-31-001",
"output_location": "gs://private-bucket/results/batch-001.json",
"registry_snapshot": {
"registry_version": "registry-20260807T0100-abc123",
"uri": "gs://private-bucket/registries/mapping_registry.registry-20260807T0100-abc123.json",
"generation": "1722980000000000",
"sha256": "..."
},
"reports": [
{
"report_id": "fax-123",
"observations": [
{"raw_name": "LDL Cholesterol", "value": "133", "unit": "mg/dL", "loinc_class_hint": "CHEM"}
]
}
]
}
Run it locally with:
python -m loinc_mapper run-worker `
--input batch.json `
--output batch-result.json `
--umls-path assets/umls/2026AA
The worker creates one mapper, calls map_many() over the flattened batch,
then restores report boundaries in the JSON result. The main project owns any
queue acknowledgement or callback delivery; it should persist the full result
before attempting Elation submission. It must download and verify the explicit
active registry snapshot before constructing the mapper; BatchWorker verifies
that the loaded registry version matches the batch reference.
Profile Matrix
Use one frozen, expert-reviewed corpus to compare the deployed path with diagnostic profiles:
python -m loinc_mapper evaluate `
--input evaluation/gold/holdout.csv `
--details evaluation/gold/profile_matrix.json `
--profile-matrix `
--umls-path assets/umls/2026AA
The report includes accepted precision, top-1/top-3, coverage, abstention, dangerous false positives, per-stage mean/p50/p95 latency, and fast-path equivalence against the full pipeline. Do not activate a change that loses a known-correct regression mapping, reduces accepted precision, or creates a dangerous false positive.
Cloud Run Job Baseline
Start with a scale-to-zero Cloud Run Job, one mapper process per task, 4 vCPU
and 8 GiB. Batch reports for up to 60 seconds or 20 reports / 500 rows before
submitting one task. Package only compact serving artifacts; mount or fetch raw
licensed UMLS data from private storage only for rebuild/audit jobs.
The pinned SapBERT files must be present in the worker image or mounted model volume before startup. The mapper uses local-only model loading so a production task never silently downloads a model from Hugging Face. On Windows development hosts, configure a sufficiently large paging file; a paging-file allocation failure is an infrastructure issue, not a reason to substitute an unvalidated smaller model.
Do not introduce Redis or ClickHouse until measurements show that multiple warm workers need a shared cache. SQLite plus FTS5, FAISS, and the in-process LRU keep this module simple enough to operate within the initial cost target.