Learning architecture
What happens to the words it can't resolve.
Every coding pipeline meets local wording it does not recognise — an abbreviation, a house style, a term from 1974. Guessing puts a wrong code in a clinical record. Discarding loses the event. This framework does neither: it retains the wording, queues it for a person, learns the mapping locally, then applies that mapping back across the histories already in the system.
This page is the inside of modules 04 to 06 — extraction, coding and review. If you have not read how it works yet, start there.
The architecture at a glance
Three flows, across the seven modules.
The framework first builds its reference layer. It then uses that layer to process patient histories. Anything it cannot confidently classify enters a learning loop instead of being guessed.
Build the reference layer
Vocabulary imports a SNOMED CT release; Anatomy cross-references the installed image catalogue against SNOMED body structures. Built once, queried locally at runtime.
Turn documents into history
Ingestion feeds a queue; Extraction interprets the source in two passes; Coding resolves concepts and anatomy against the reference layer.
Resolve what was unknown
Unclassified terminology enters the review queue. Once a person teaches the mapping, existing histories are re-indexed without repeating the original extraction.
Flow 1
Build the reference layer.
This happens once, before any patient history is processed. It creates the local terminology and anatomy lookups the rest of the pipeline queries — no terminology service calls at runtime.
SNOMED CT + anatomical image build
The RF2 release and the registered image catalogue are combined into a local clinical-to-visual lookup layer.
SNOMED RF2 release
The Snapshot release is placed in the configured private server folder.
configured release pathSNOMED import
Concepts, descriptions and relationships are parsed, validated and indexed locally.
concepts · terms · hierarchyClinical terminology foundation
Fast concept and term lookups are built for the patient-history pipeline.
local lookup, not remote terminology callsBuild Image DB
The installed anatomical layers are matched against SNOMED body structures, synonyms and hierarchy.
concept ↔ anatomy ↔ PNG layerClinical + visual knowledge base
The system now knows both what a clinical concept means and where it can be represented on the body.
terminology + anatomy + layer mappingsFlow 2
Import a patient history from file or API.
Both routes create the same durable document record and feed the same queue, so a bulk backfill of twenty years of correspondence and a single letter take the same path.
Patient history intake
TXT upload
An operator uploads a raw text history through the browser.
API import
An external system submits raw history data programmatically.
Patient + source document
The source text, patient link and provenance are stored before extraction starts.
Ingest queue
Persistent workers claim documents and process multiple patients concurrently.
Two-pass clinical interpretation
The language model interprets unstructured prose. It does not assign codes — that is the coding module, working against the local terminology store.
Understand the document structure
Durable source windows are divided into coherent clinical sections or encounters. A dated presentation, its issues and its treatment remain together rather than being split by arbitrary text boundaries.
Extract structured clinical facts
Diagnoses, injuries, procedures, investigations, medications, treatments, follow-up and status changes are extracted with dates, evidence and provenance.
From extracted facts to a patient timeline
Document-wide reconciliation
Copied problem lists, duplicate mentions, status changes and non-events are reconciled across the full source.
SNOMED resolution
Local indexes retrieve and resolve appropriate SNOMED concepts for the extracted clinical facts.
Anatomical mapping
SNOMED/body-structure relationships are resolved to the registered image layers where possible.
Clinical audit
Coverage, event fidelity, duplicate handling, terminology resolution and timeline consistency are checked.
Longitudinal patient history
Safe events are committed as dated, coded, anatomically linked records with source evidence retained.
Flow 3
The human-guided learning loop.
Unknown terminology is never silently forced into a category. It is retained with its original wording, reviewed by a person, learned locally, and then applied back to the data already in the system. This is the part that makes a twenty-year backfill improve rather than fossilise.
Unresolved → learned → re-indexed
Unclassified event
A clinical fact cannot be confidently mapped to the current local terminology/anatomy knowledge.
Learning queue
The original wording, proposal and context are retained for review rather than discarded.
Teach the local mapping
Reviewers connect local or unusual wording to the appropriate SNOMED concept and anatomy layer where applicable.
Re-index patient history
Previously unresolved events are run through the improved local mappings without repeating the original document extraction.
What actually learns?
The distinction matters.
The standard does not change
The imported SNOMED release remains the terminology authority. The framework does not invent new SNOMED concepts or rewrite the standard.
Mappings improve
The system can learn local aliases, historical wording, abbreviations and concept-to-anatomy/image relationships that were previously unresolved.
Uncertainty stays visible
Anything unresolved stays reviewable and traceable to its source. The loop is controlled by the organisation running it. The framework makes no clinical decision at any point in it.
The whole framework
One end-to-end view.
The architectural idea
AI reads. SNOMED grounds. Humans govern. The system remembers.
The language model is used where it is strongest: interpreting messy, unstructured clinical prose. SNOMED CT and the local lookup store provide the terminology foundation. Human review handles uncertainty and teaches local mappings. Re-indexing then applies that knowledge back across existing patient histories.
The result is not simply an AI summary. It is a traceable clinical classification pipeline with a visual anatomy model and an explicit improvement loop.
See the document pipeline
Follow one unstructured clinical document from source text to coded, anatomically located events.
The learning queue is where the real work is
Reviewing the resolution strategy is one of the six things this project needs most. Founding implementers get repo access and a say in it.