Skip to content

Performance, Profiles, and Worker Batches

The mapper is intentionally safety-first. Performance work removes duplicated computation; it never removes the clinical checks that decide whether a LOINC code can be emitted.

Production Flow

For every observation that is not a verified exact universal-template fast-path case:

normalize
-> AxisFactExtractor + UniversalSemanticIndex
-> deterministic SQLite retrieval
-> ScispaCy mention and abbreviation processing
-> local UMLS CUI retrieval
-> FAISS/SapBERT semantic retrieval
-> UCUM and six-axis validation
-> feature ranking
-> confidence and margin gate

The only permitted shortcut is an exact reviewed universal_template alias with review_status=verified_active and fast_path_enabled=true. It still proves:

  • the target code is active in the pinned LOINC release;
  • declared unit, class, property, method, and specimen requirements are met;
  • UCUM/unit and six-axis validation pass;
  • no inferred context conflicts exist; and
  • the confidence policy accepts the single reviewed target.

Fast-path provenance lists skipped semantic stages. A contained alias, legacy registry row, source-specific evidence record, unknown unit, or missing required specimen/method always runs the full pipeline.

Settings

Copy .env.example to the ignored .env file and keep production settings enabled:

PIPELINE_PROFILE=production
IS_UMLS_ON=true
IS_FAISS_ON=true
IS_SAPBERT_ON=true
IS_VALIDATION_ON=true
IS_CONFIDENCE_ON=true
IS_REGISTRY_FAST_PATH_ON=true
MAPPER_CPU_THREADS=4

production rejects any disabled evidence stage. evaluation is the only profile that allows controlled ablations. Any evaluation result is marked execution.experimental=true, and ElationAdapter refuses to create a payload from it.

Validation and confidence cannot be disabled in any profile. Do not use a profile experiment as a deployment setting.

Why It Is Faster

  • The SQLite catalog loads once per worker. Class/type membership and token statistics are kept in memory for O(1) scope checks.
  • The release-derived universal component/axis index is compiled once from that loaded catalog. It creates source-neutral candidates from facts such as urine, CSF, 24-hour duration, and stated methods; it does not query a per-laboratory mapping table for ordinary rows.
  • Deterministic lookup uses the indexed alias_tokens table first. FTS5 is a bounded prefix fallback, not a broad query over every matching alias.
  • Character recovery uses distributed 3-of-4 trigram intersections. It still tolerates a single damaged OCR trigram, but avoids the old broad trigram OR scan for words such as cholesterol.
  • Deterministic candidates are generated once per verbatim row name within a batch (the evidence depends on the exact string), and context extraction is cached per normalized name. Text normalization, unit parsing and the semantic query variants are memoized process-wide: they are pure functions of the pinned vocabularies, and the validator calls them about ten times per candidate.
  • Alias similarity is gated by SequenceMatcher's two upper bounds (normalization.best_alias_ratio, ratio_if_at_least): the exact ratio is computed only when it can still beat the best score or reach the threshold, and ties resolve exactly as the ungated scan did.
  • ScispaCy processes a report through nlp.pipe; parser/tagger components not needed for linking are disabled.
  • The UMLS linker owns one read-only SQLite connection and caches mention-to-CUI results for the worker lifetime; exact lookups and the fuzzy decision stay per mention, while the CUI-to-LOINC bridge is queried once per batch of mentions.
  • The result cache is keyed on the routing projection of report_context (resolve_routing is its only mapping-path reader), so identical rows of different reports share a cached answer while a job id or filename never changes one; cached results are stored as pickled bytes and the runtime manifest is copied by JSON parse, both value-equal and isolated from caller mutation.
  • A linked UMLS name receives bounded exact/token local expansion rather than replaying the complete OCR deterministic pipeline once for every sibling CUI. Direct CUI-to-LOINC bridge candidates are always retained.
  • SapBERT encodes unique queries in batches. FAISS searches are batched too.
  • Code-level SapBERT vectors avoid re-embedding candidate descriptions during reranking.
  • Optional class shards search an official LOINC class first, then use a laboratory/global fallback if a shard produces no scoped hit.

Every result records provenance.stages.timings_ms, candidate counts, profile, artifact versions, cache hits, and fallback evidence. Inspect those fields before changing a threshold.

timings_ms covers only the per-row stage (_map_one). The batch prologue that map_many runs once per report is reported separately, with the same values copied onto every result of that batch:

  • provenance.stages.batch_timings_ms — panel_resolve_ms, row_probes_ms (cache, force-exact and fast-path probes), deterministic_batch_ms, umls_link_batch_ms (ScispaCy nlp.pipe plus per-mention UMLS lookups), faiss_prefetch_ms (query encoding plus the batched FAISS search), map_one_rows_ms, total_ms.
  • provenance.stages.batch_counters — the delta of process-lifetime counters for that batch: result_cache_hits, force_exact_results, fast_path_results, pending_rows, structure_invalid_rows, umls_mentions_linked, umls_mention_cache_hits, umls_fuzzy_fallbacks (+ umls_fuzzy_ms), umls_trigram_fallbacks (+ umls_trigram_ms), semantic_cache_hits, semantic_cache_misses, catalog_fuzzy_queries (+ catalog_fuzzy_ms), embedding_cache_hits, embedding_encoded (+ encode_ms), faiss_queries (+ faiss_search_ms).
  • provenance.stages.batch_size.

When a report's wall clock is far above the sum of its rows' total_ms, the difference is in batch_timings_ms; the counters say which fallback paid.

Measuring a change

Measured on 2026-09-17 with the production snapshot and force overrides stripped: the cold pass fell from 260.1 s to 132.2 s for the 55 rows after the results-identical round (memoized normalization and unit parsing, gated alias similarity, batched UMLS bridge, routing-keyed result cache); validation_ms mean 712 -> 69 (p95 1748 -> 169), umls_candidate_union_ms mean 4320 -> 2448, deterministic_batch_ms 11.9 -> 8.4 s. The remaining cold cost is the per-concept generate_semantic misses inside umls_candidate_union_ms, the next target. Pass 2 of the benchmark clears the result cache first (the routing-keyed cache would otherwise serve every row of a repeated fixture in well under a second, which is the production case for recurring rows but not a pipeline measurement); --keep-result-cache shows the cache-read case instead.

tools/benchmark_map_many.py is the operator gate for performance work. It maps the 55-row de-identified fixture (tests/fixtures/ocr_recovery_review_cases.20260910.json) through build_mapper on the real artifact bundle, strips the force-exact overrides by default so the certified rows exercise the full pipeline, and prints the batch timings, counters, and per-stage p50/p95 for each pass:

$env:PYTHONPATH = "$PWD\src"
python tools\benchmark_map_many.py --umls-path assets/umls/2026AA --repeat 2 `
  --write-results evaluation/runs/baseline.json
# after the change
python tools\benchmark_map_many.py --umls-path assets/umls/2026AA --compare evaluation/runs/baseline.json

--compare exits 1 when any row's mapping identity (status, code, ordered candidates, confidence, margin, reasons, axis and unit evidence, diagnostics) differs from the saved run. Pass two shows warm-cache behaviour. Results go under the gitignored evaluation/runs/; the tool writes nothing under config/. On a Windows host, a SapBERT OSError 1455 (paging file too small) is an infrastructure limit, not a reason to substitute a smaller model.

Historical note: the single-row deterministic retrieval of VLDL Cholesterol Cal measured 611.6 ms after the alias-index fix (down from 40-48 s). End-to-end numbers come from the benchmark above and from the production batch_timings_ms.

The raw 2026AA fallback now performs intersection-based token lookup and measured 5.3 s for that same UMLS name, down from 17.3 s; it is still too slow for the report target. The compact serving artifact remains a required deployment gate for production latency measurements.

UMLS Serving Artifact

The raw UMLS umls.sqlite3 may be large because it contains broad Metathesaurus and older trigram structures. Build this compact runtime artifact once after verifying raw assets:

python -m loinc_mapper build-umls-serving `
  --asset-path assets/umls/2026AA `
  --catalog assets/loinc/2.82/catalog_v5.sqlite3 `
  --output assets/umls/2026AA/serving/umls_loinc.sqlite3

It retains UMLS aliases only when their CUI bridges to an active local LOINC code and creates exact, FTS5, and trigram indexes. The source release is validated by checksum during this build. At worker startup, the serving manifest validates release, artifact shape, and file size without rehashing the raw 14 GB file. Runtime provenance should then show serving_exact+fts5+trigram.

Keep raw licensed data private and immutable so the serving artifact can be rebuilt and audited later.

For the background PowerShell build, follow progress on standard error:

Get-Content evaluation\logs\umls_serving_build.err.log -Wait

FAISS Derived Artifacts

An existing terms.faiss can gain code-level vectors without a new SapBERT encoding run:

python -m loinc_mapper build-code-vectors `
  --index assets/loinc/2.82/terms.faiss `
  --metadata assets/loinc/2.82/terms.json

Optional laboratory/class shards are also derived from the same global index:

python -m loinc_mapper build-vector-shards `
  --catalog assets/loinc/2.82/catalog_v5.sqlite3 `
  --index assets/loinc/2.82/terms.faiss `
  --metadata assets/loinc/2.82/terms.json `
  --output assets/loinc/2.82/shards

Shards are a latency optimization only. SQLite remains the authoritative LOINC catalog, and every shard result still undergoes the same laboratory-type, UCUM, and six-axis safety checks. Class agreement is advisory evidence; a laboratory-wide fallback prevents a mixed report section from erasing a valid candidate. Benchmark class shards before shipping them because they consume additional disk space.

Batch Worker Contract

The main fax project should submit up to 20 reports or 500 observations to one long-lived worker process. MappingBatch version 1 looks like this:

{
  "schema_version": "1",
  "batch_id": "fax-batch-2026-07-31-001",
  "output_location": "gs://private-bucket/results/batch-001.json",
  "registry_snapshot": {
    "registry_version": "registry-20260807T0100-abc123",
    "uri": "gs://private-bucket/registries/mapping_registry.registry-20260807T0100-abc123.json",
    "generation": "1722980000000000",
    "sha256": "..."
  },
  "reports": [
    {
      "report_id": "fax-123",
      "observations": [
        {"raw_name": "LDL Cholesterol", "value": "133", "unit": "mg/dL", "loinc_class_hint": "CHEM"}
      ]
    }
  ]
}

Run it locally with:

python -m loinc_mapper run-worker `
  --input batch.json `
  --output batch-result.json `
  --umls-path assets/umls/2026AA

The worker creates one mapper, calls map_many() over the flattened batch, then restores report boundaries in the JSON result. The main project owns any queue acknowledgement or callback delivery; it should persist the full result before attempting Elation submission. It must download and verify the explicit active registry snapshot before constructing the mapper; BatchWorker verifies that the loaded registry version matches the batch reference.

Profile Matrix

Use one frozen, expert-reviewed corpus to compare the deployed path with diagnostic profiles:

python -m loinc_mapper evaluate `
  --input evaluation/gold/holdout.csv `
  --details evaluation/gold/profile_matrix.json `
  --profile-matrix `
  --umls-path assets/umls/2026AA

The report includes accepted precision, top-1/top-3, coverage, abstention, dangerous false positives, per-stage mean/p50/p95 latency, and fast-path equivalence against the full pipeline. Do not activate a change that loses a known-correct regression mapping, reduces accepted precision, or creates a dangerous false positive.

Cloud Run Job Baseline

Start with a scale-to-zero Cloud Run Job, one mapper process per task, 4 vCPU and 8 GiB. Batch reports for up to 60 seconds or 20 reports / 500 rows before submitting one task. Package only compact serving artifacts; mount or fetch raw licensed UMLS data from private storage only for rebuild/audit jobs.

The pinned SapBERT files must be present in the worker image or mounted model volume before startup. The mapper uses local-only model loading so a production task never silently downloads a model from Hugging Face. On Windows development hosts, configure a sufficiently large paging file; a paging-file allocation failure is an infrastructure issue, not a reason to substitute an unvalidated smaller model.

Do not introduce Redis or ClickHouse until measurements show that multiple warm workers need a shared cache. SQLite plus FTS5, FAISS, and the in-process LRU keep this module simple enough to operate within the initial cost target.