Seven modules, each with its own inputs, outputs and storage. Every one is a seam: run it as shipped, or replace it with your own and keep the rest. These are the names used everywhere — on this site, in the repo and in the code.
Two modules are built once and produce the reference layer. Five run for every record. The anatomical view and the coded data fall out of the same pass; what you then do with that coded data is your own build.
01Vocabulary
Imports a SNOMED CT release into a queryable local structure. It runs in resumable stages, each checked on input hash and row count, so a failed import restarts where it stopped rather than from the beginning.
Input
A SNOMED CT RF2 Snapshot release, under your own affiliate licence. No terminology content ships with the framework.
Process
Concepts, descriptions, relationships and language refsets loaded in turn, the release verified, then the derived lookups built.
Output
A local terminology store the rest of the pipeline queries directly — no terminology service calls at runtime.
Extend
Swap the edition or release. National extensions and local refsets load through the same path.
02Anatomy
Bridges the terminology to the picture. SNOMED CT body structures are walked into a hierarchy and mapped to the drawable regions of the figure, so an anatomical concept can resolve to a place on the body.
Input
The imported body-structure hierarchy, plus the image layer files themselves.
Process
Hierarchy built, concepts mapped to layers, then audited against the image directory — the files are the source of truth and the rows are annotations on them.
Output
A concept-to-region lookup, with a broader anatomical fallback for concepts that resolve only to a general area.
Extend
Bring your own figure. Replace the image set, re-run the mapping, and nothing upstream changes.
02aThe layer library
The anatomy module is only as good as the artwork underneath it, and this is the part that took longest. Two complete figures — male and female — built as separate registered layers rather than as pictures, so any combination of events can be lit at once without the body being redrawn and without the anatomy drifting between renders.
Worth separating two things here. The structure — the layer registration, the filename contract, explicit laterality — is the durable part and it works. The draughtsmanship is mine, and I am not an illustrator. Replacing the artwork without touching anything upstream is a deliberate property of the design, and it is one of the six things this project is asking for.
organ_gallbladder_right_front.png
systemregionlateralityview
The filename is the contract. Every layer declares its own system, region, laterality and view, so a file can be matched without a lookup table telling you what it contains — the directory is the source of truth and the database rows are annotations attached to it. Laterality is explicit (left, right, bilateral, midline) because SNOMED CT models laterality explicitly too, and a left knee is not a right knee.
baseThe skeletal figure everything else registers against, front and back
jointShoulder, elbow, wrist, hip, knee, ankle, each side separately
muscleMuscle groups by region, including the posterior view
organViscera, glands and sex-specific organs on both figures
vascularCoronary and limb vasculature
lymphaticNodal regions — head, chest, abdomen, inguinal, limbs, spine
skinSurface regions, so integumentary findings have somewhere to land
systemFunctional systems that cross regions — cardiac conduction, upper airway, reproductive
Scale
Close to three hundred layers across the two figures, anterior and posterior, hand-assembled rather than generated. Coverage is honest but incomplete — rarer structures fall back to a broader region.
Why layers
A single generated image degrades every time you ask for a different combination of highlights. Fixed base plus composited layers gives perfect registration and repeatable output.
Resolution
28231008 |Gallbladder structure| → organ_gallbladder_right_front.png — with a broader regional layer as fallback when a concept resolves only to a general area.
03Ingestion
Takes documents as they actually exist. Discharge summaries, operation notes and specialist letters go in unmodified — no templates, no forms, no pre-structuring. Intake is a webhook onto a queue, so the flow of health data into the system stays under your control and on your schedule.
Input
Raw documents and text, posted to the intake webhook or uploaded directly.
Process
Work lands on a queue, then text is extracted, split into workable chunks and processed as a job that survives a restart.
Output
Chunked source text with its provenance intact, ready for extraction.
Extend
Wire the webhook into whatever already holds your correspondence. Rate, batching and what qualifies to be sent are yours to set — nothing pulls data on its own.
04Extraction
Reads the source in two passes — segment the document, then extract events from those segments — and proposes discrete clinical events, each carrying the passage it came from. Nothing here reaches a patient record.
Input
Chunked source text.
Process
Pass one divides the text into coherent clinical sections; pass two extracts dated events from each. Both run against the configured language model, within a request and token budget you set.
Output
Proposed events with a date, a summary and the source evidence behind them.
Extend
The model provider is configuration, not architecture. Swap it, run it locally, or replace this module outright — the contract is text in, proposed events out.
05Coding
Binds each proposed event to SNOMED CT and resolves its anatomy where the terminology supports it. This is the module the project exists for.
Input
Proposed events.
Process
Clinical text matched against the pre-built term lookups to a SNOMED CT concept; anatomy resolved through SNOMED CT relationships where they are available.
Output
Coded, dated events, with an anatomical region where one can be determined and a flag where it cannot.
Extend
Unmatched concepts and ambiguous anatomy route to the learning queue rather than being guessed at — see the learning architecture. That queue is where local terminology work happens.
06Review
The gate, wherever you choose to set it. Events arrive proposed, scored and shown with their evidence; how much passes automatically and how much a person sees is policy, and the policy is yours.
Input
Coded events with their source evidence and confidence signals.
Process
Each run scored on coverage, terminology resolution and event fidelity, with blockers and warnings surfaced rather than buried.
Output
An approved, corrected or rejected decision per event, with an audit trail of which were accepted automatically and which a person touched.
Extend
Tighten or loosen what passes without a human. A pilot on historic records and a live clinical deployment are the same code with different thresholds.
07Patient history
What a clinician actually opens: timeline, body map and events reading as one thing.
Input
Approved, coded events.
Process
Events placed on the timeline and composited onto the anatomical figure at the selected point in time.
Output
A history readable in seconds — and the same events, coded and dated, sitting in your own database ready for whatever you build next.
Extend
This is the layer most worth replacing with your own. The coded history underneath it is the durable part.
Verify before you plan against this. These descriptions track a live implementation and will move as it hardens. If you're evaluating seriously, ask — I'd rather walk you through the real thing, in the repo, than have you build on a web page.
Want to shape this?
Founding implementers get an invitation to the private repo, and a say in which module gets hardened first.