XyDromatics™ De-Identification Engine —
raw PHI in, research-safe data out.
The one-way de-identification appliance. Accept studies and
reports across an untrusted network over mutually-authenticated
encrypted channels, apply HIPAA Expert-Determination
de-identification — pseudonymize identifiers, interval-shift
dates, scrub free text and burned-in pixels — and forward
de-identified output, plus PHI-free coded concepts, to VNA
Research. The trust boundary between clinical data and research.
Runs Synthology-hosted in the cloud or customer-operated
on-prem — the same trust boundary either way. The
re-identification key lives only on the engine and never
crosses the wire.
One job, done thoroughly: PHI in, research-safe data out.
Secure cross-network ingest, Expert-Determination
de-identification, a content gate that reaches into free text
and burned-in pixels, HL7 report de-identification, and a
PHI-free coded-concept harvest for research cohorts.
Secure cross-network ingest
Accept studies and reports pushed across an untrusted network, from any source, over an encrypted, mutually-authenticated channel. Every channel is fail-closed — an unauthenticated or unpinned peer is rejected before a single object is read.
DICOM-TLS C-STORE — secure DIMSE for C-STORE-only modalities and legacy PACS (TLS 1.2/1.3)
MLLP-over-TLS — secure HL7 v2 ORU for EMR / RIS report sources
DICOMweb STOW-RS over HTTPS — for cloud and third-party PACS sources
Mutual TLS with SHA-256 client-certificate thumbprint pinning on every channel
Optional sealed bearer token as an alternative to mTLS on the DICOMweb channel
Off by default — on-prem deployments keep plain DIMSE / MLLP inside the trusted network
De-identification core
The de-identification itself, applied under the HIPAA Expert-Determination method. Identity is replaced deterministically so the same patient stays linkable across studies without any correlation table.
Keyed-HMAC pseudonymization of patient and study identifiers (per-site key, never leaves the engine)
Interval-preserving per-patient date shift (DICOM Attribute Modification 113107) — exam intervals and chronological order preserved, absolute dates obscured
VR-aware tag substitution — replacements always fit the target value representation (no silent blanking)
Private / proprietary tag removal (default on)
Deterministic study-UID remapping — the de-identified study carries a stable, re-derivable UID
Optional operator-chosen sequential labeling (licensed; requires a captured §164.514(c) attestation)
Content gate — free text, pixels, documents
Identifiers hide in more than structured tags. The content gate finds and removes PHI in narrative reports, burned-in image pixels, and encapsulated documents — or quarantines what it cannot safely clean.
Free-text SR narrative scrubbing (shared PHI heuristic scrubber)
Burned-in pixel OCR redaction — local Tesseract renders the pixels and blacks out detected text (never cloud OCR)
Encapsulated-document handling (PDF rasterize + OCR-redact + rebuild, or text scrub)
Per-lane policy: clean, quarantine, or reject — operator-configurable per content type
Quarantined objects are sealed on the engine for operator review — never silently forwarded
HL7 v2 ORU de-identification
Reports de-identify with the SAME per-patient pseudonym and date-shift offset as the images, so a de-identified report lands on its de-identified study's timeline with no correlation table and in any arrival order.
Cross-protocol demographic alignment — HL7 and DICOM derive the pseudonym + date shift from the same MRN + site key
OBX narrative scrubbed via the shared free-text PHI scrubber
Store-and-forward with fail-closed NACK — the sender holds any message the engine could not fully de-identify and forward
Forwards the de-identified report to a downstream HL7 research sink
Coded-concept harvest → research cohorts
Beyond de-identified images, the engine can harvest standards-coded clinical concepts from report-bearing objects into a PHI-free, pseudonym-keyed cohort store — the input to a Clinical Cohort Builder. Opt-in and gated.
PHI-free coded-concept batches keyed by the study pseudonym
Two-person export control — a licensed add-on (deid_cohort_export) AND a current §164.514(b)(1) Expert-Determination attestation
k-anonymity floor enforced on cohort export
Fail-open harvest: a harvest or ingest error never affects the de-identified image forward
Operations, durability & administration
A production appliance: durable store-and-forward, sealed configuration and secrets, PHI-free health telemetry, and the shared administration surface used across the fleet.
Durable outbound spool — store-then-drain so a downstream outage or restart never loses an instance
Quarantine + dead-letter queues with operator review
Whole-configuration encryption at rest (SYNTHIMG envelope; reversible secrets sealed, never emitted)
Tamper-evident, hash-chained HIPAA audit log with configurable retention
Web admin UI with shared sidebar shell + role-based access control
PHI-free /api/health telemetry (queue depth, disk, channel posture) for pool + gateway monitoring
Dataflow
The engine is the trust boundary.
Raw PHI arrives from the customer site over an encrypted,
mutually-authenticated channel — DICOM-TLS C-STORE,
MLLP-over-TLS, or DICOMweb STOW-RS. The engine de-identifies at
the boundary and forwards only de-identified output to VNA
Research. It occupies the same position whether Synthology hosts
it in the cloud or the customer runs it on-prem.
Use cases
Five deployment patterns.
Cloud-hosted or on-prem, image or report, external-vendor
shielding or research-cohort building — the engine is the
single de-identification point for all of them.
Pattern 1
Synthology-hosted research ingest (cloud)
A customer sends studies to a Synthology-hosted De-Identification Engine across the internet. Raw PHI crosses only inside a mutually-authenticated, encrypted channel; nothing re-identifiable is ever stored in the cloud.
Flow
1Customer PACS / modality sends over DICOM-TLS C-STORE (or DICOMweb STOW-RS) with a pinned client cert
2The engine authenticates the peer, decrypts in memory, and de-identifies at the trust boundary
3Identifiers pseudonymized, dates interval-shifted, free text + burned-in pixels scrubbed
4De-identified study forwarded to VNA Research; PHI-free coded concepts harvested to the cohort store
5No PHI and no re-identification key persist in the cloud — only de-identified output
Pattern 2
On-prem de-identification appliance
The customer runs the whole chain inside their own network. The engine is the internal gate between clinical archives and a research environment — TLS optional, since traffic never leaves the trusted network.
Flow
1Clinical PACS / Router forwards studies to the on-prem engine over plain C-STORE (inside the trusted LAN)
2The engine de-identifies under the site's Expert-Determination policy
3De-identified output lands in an on-prem VNA Research archive
4Coded concepts populate the local Clinical Cohort Builder
5The re-identification key stays on the customer's engine; researchers see only de-identified data
Pattern 3
AI-vendor / external-reader PHI shielding
Send de-identified studies to an AI vendor or outside reading service that is not in BAA scope. They receive only pseudonymized data; the mapping stays on the engine.
Flow
1Source study de-identified with keyed-HMAC pseudonymization + date shift
2De-identified study sent to the AI vendor / external reader
3The vendor analyzes and returns findings tagged with the same pseudonyms
4The re-identification mapping — held only on the engine — links findings back to the right patient
5The vendor never received PHI; the clinical workflow still benefits from the findings
Pattern 4
Research report de-identification (HL7)
ORU reports from the EMR / RIS de-identify alongside the images and land on the matching de-identified study's timeline, so a research reporting environment has both the pixels and the narrative — with no PHI.
Flow
1EMR / RIS emits ORU^R01 over MLLP-over-TLS
2The engine de-identifies PID / OBR / OBX / PV1 / ORC and scrubs the OBX narrative
3HL7 and DICOM derive the same pseudonym + date shift from the same MRN + site key
4The de-identified report is store-and-forwarded to the research HL7 sink (fail-closed on any failure)
5Report and study align on the timeline with no correlation table
Pattern 5
Clinical-cohort building for research
A research program needs cohorts defined by coded clinical concepts, not free text. The engine harvests standards-coded concepts into a PHI-free, pseudonym-keyed store that a Clinical Cohort Builder queries under k-anonymity.
Flow
1Harvest lane extracts standards-coded concepts from pre-scrub report objects (opt-in, licensed)
2PHI-free coded batches keyed by the study pseudonym are posted to VNA Research
4Export is gated: the deid_cohort_export add-on AND a current §164.514(b)(1) Expert-Determination attestation, both required
5Cohorts export under an enforced k-anonymity floor
Integration points
Sources in, de-identified output out.
Sources (senders)
Modalities + PACS over DICOM C-STORE (plain on-prem, or DICOM-TLS across a network)
EMR / RIS over HL7 v2 ORU (plain MLLP, or MLLP-over-TLS)
Third-party PACS / cloud over DICOMweb STOW-RS (mTLS or sealed bearer)
XyDromatics Router — can route studies to the engine for de-identification
Destinations (de-identified output)
XyDromatics VNA Research — de-identified DICOM archive (Research-Use-Only)
VNA Research coded-concept cohort store (PHI-free) — the Clinical Cohort Builder input
Downstream HL7 research reporting sink (de-identified ORU)
AI vendors / external reading services — pseudonymized studies only
Cross-product orchestration
Shared per-site encryption key — consistent pseudonyms across Router, VNA family, and the engine
Shared DICOM SCP library — the same C-STORE receive path used fleet-wide (now with optional DICOM-TLS)
Unified key-escrow (IKeyEscrow) — every reversible secret sealed under one customer master
Synthology QMS — change control on de-identification policy
Identity, audit, observability
OIDC / SAML + Active Directory / Azure AD for the web admin UI
Hash-chained, tamper-evident HIPAA audit log
PHI-free health telemetry (SynthIQ pool + SynthGateway console)
Per-peer, per-channel audit of every accepted association
System requirements & sizing
Sized to study volume + content-gate depth.
Pixel-OCR redaction and encapsulated-document handling are the
heaviest operations; sizing reflects study volume and how much
of the content gate is enabled.
Tier
Study volume
CPU
RAM
Storage
Typical site
Small (single facility)
< 5,000 studies / day
4 vCPU
8 GB
250 GB (spool + quarantine + audit)
Imaging center feeding a single research program.
Medium (community hospital)
5,000 – 25,000 studies / day
8 vCPU
16 GB
1 TB
Full modality mix with pixel-OCR + HL7 report de-id.
Everything, plus the coded-concept harvest lane and the Clinical Cohort Builder — under two-person Expert-Determination export control.
Everything in Secure Ingest & Content Gate
Coded-concept harvest → PHI-free pseudonym-keyed cohort store
deid_cohort_export add-on with §164.514(b)(1) attestation gate
k-anonymity-floored cohort export
Active-active HA pair · premium support eligible
Per-deployment licensing. Multi-product bundles (De-ID Engine +
VNA Research + Router) get bundle pricing and shared site-key
provisioning. The Synthology-hosted deployment is offered as a
managed service — see
Managed services.
Documentation
The controlled-document set.
Current GA release
v1.0.1.63·released 2026-08-05
Every De-Identification Engine deployment ships with the documents below, all managed under SynthQMS document control. The DICOM Conformance Statement is publicly available; the rest are shared under mutual NDA.
What’s new·v1.0.1.63
Raw PHI in, research-safe data out — a dedicated de-identification pipeline
▸Secure, fail-closed ingest over mutually-authenticated TLS — DICOM C-STORE, HL7 v2, and DICOMweb STOW-RS — so an unrecognized peer is turned away before any object is read.
▸De-identification designed to support a HIPAA Expert-Determination workflow — identifiers are pseudonymized deterministically, so the same patient stays linkable across studies, and dates are shifted per patient to preserve intervals while obscuring the calendar.
▸A content gate beyond structured tags — scrubbing free-text report narratives, redacting burned-in pixel text with on-box OCR (never the cloud), and handling encapsulated PDFs — quarantining anything it cannot safely clean instead of forwarding it.
▸De-identify HL7 v2 result reports onto the same de-identified study’s timeline, and optionally harvest standards-coded clinical concepts into a PHI-free, pseudonym-keyed cohort store — gated by two-person export control and a k-anonymity floor.
▸Runs Synthology-hosted or on-prem, with the re-identification key kept only on the engine, durable store-and-forward, secrets sealed at rest, and a tamper-evident HIPAA audit log. Ships with the approved EULA and clears a full QA + security-test pass.
Annex to the Synthology Master EULA (DOC-2026-063)
DOC-2026-450
Penetration Test Results (De-Identification Engine)
Available under NDA
DOC-2026-449
QA Runner Manual (De-Identification Engine)
Internal QA + customer-validated test runs
DOC-2026-447
User Guide (De-Identification Engine)
Operator + administrator workflows
DOC-2026-448
Installation Guide (De-Identification Engine)
Includes MSI deploy + Linux tarball install
Frequently asked
The questions prospects ask.
How is this different from anonymization inside the Router or the VNA?
The Router and the VNA family can pseudonymize on export as part of a routing or archive workflow. The De-Identification Engine is the dedicated, one-way appliance whose whole job is turning raw PHI into a research-safe data set under the HIPAA Expert-Determination method — with the secure cross-network ingress, the content gate (free text + burned-in pixels + documents), HL7 report de-identification, and the coded-concept cohort harvest that a general router does not carry. It is the trust boundary between clinical data and a research environment.
Do you support secure ingest over both DIMSE and DICOMweb?
Yes — both, plus HL7. A source can push over DICOM-TLS C-STORE (secure DIMSE, for C-STORE-only modalities and legacy PACS), MLLP-over-TLS (secure HL7 v2), or DICOMweb STOW-RS over HTTPS. Every channel is mutually authenticated with SHA-256 client-certificate thumbprint pinning and is fail-closed: an unauthenticated or unpinned peer is rejected before any object is read. All of it is off by default, so an on-prem deployment inside a trusted network keeps plain DIMSE / MLLP.
Who holds the re-identification key?
Only the engine. Pseudonymization uses a keyed HMAC with a per-site key that never leaves the engine, sealed at rest under the unified key-escrow master. De-identified output carries no key and cannot be re-identified by any downstream — VNA Research, an AI vendor, or an external reader. In the Synthology-hosted deployment, no PHI and no key persist in the cloud; only de-identified output does.
Does the date shift break longitudinal research?
No — that is the point of it. Each patient gets a stable, per-patient offset (HMAC-derived), applied uniformly to every date. Exam intervals and chronological order are preserved exactly, so longitudinal analysis still works; only the absolute calendar dates are obscured. The engine stamps DICOM Attribute Modification 113107 whenever a date was actually shifted. Ages over 89 are handled per Safe-Harbor expectations.
How do reports stay aligned with their de-identified studies?
Cross-protocol demographic alignment. Both the HL7 and the DICOM paths derive the pseudonym AND the date-shift offset from the same MRN + site key, so a de-identified report lands on its de-identified study's timeline with no correlation table and regardless of which arrives first. The HL7 path is store-and-forward and fail-closed: the sender holds any message the engine could not fully de-identify and forward.
What is the coded-concept cohort harvest, and how is it controlled?
An opt-in lane that extracts standards-coded clinical concepts from report-bearing objects into a PHI-free, pseudonym-keyed store — the input to a Clinical Cohort Builder. It never stores free-text PHI blobs. Turning it on and exporting cohorts is a two-person control: it requires the deid_cohort_export licensed add-on AND a current §164.514(b)(1) Expert-Determination attestation on file, and cohort export is held to an enforced k-anonymity floor.
What is the regulatory classification?
Non-device software under FD&C Act §520(o)(1)(D) — the 21st Century Cures Act carve-out for software that transfers, stores, converts, or displays device data without interpreting it. The engine does not interpret images or reports; it de-identifies and forwards. Its de-identified output is intended for Research Use Only, and the de-identification follows the HIPAA Expert-Determination method (45 CFR §164.514(b)(1)).
Is it cross-platform?
Yes — Windows (MSI installer) and Linux (self-contained tarball), same as the rest of the fleet. The Synthology-hosted deployment runs on Linux; on-prem customers run either.
Build your de-identified research pipeline.
Cross-network ingest to a Synthology-hosted engine, or an
on-prem appliance inside your network? Tell us where your PHI
starts and where your research data needs to land, and
we’ll come back within one business day with a proposed
architecture.