XyDromatics De-Identification Engine Specialized engine · De-identification

XyDromatics™ De-Identification Engine — raw PHI in, research-safe data out.

The one-way de-identification appliance. Accept studies and reports across an untrusted network over mutually-authenticated encrypted channels, apply HIPAA Expert-Determination de-identification — pseudonymize identifiers, interval-shift dates, scrub free text and burned-in pixels — and forward de-identified output, plus PHI-free coded concepts, to VNA Research. The trust boundary between clinical data and research.

Runs Synthology-hosted in the cloud or customer-operated on-prem — the same trust boundary either way. The re-identification key lives only on the engine and never crosses the wire.

Capabilities

One job, done thoroughly: PHI in, research-safe data out.

Secure cross-network ingest, Expert-Determination de-identification, a content gate that reaches into free text and burned-in pixels, HL7 report de-identification, and a PHI-free coded-concept harvest for research cohorts.

Secure cross-network ingest

Accept studies and reports pushed across an untrusted network, from any source, over an encrypted, mutually-authenticated channel. Every channel is fail-closed — an unauthenticated or unpinned peer is rejected before a single object is read.

  • DICOM-TLS C-STORE — secure DIMSE for C-STORE-only modalities and legacy PACS (TLS 1.2/1.3)
  • MLLP-over-TLS — secure HL7 v2 ORU for EMR / RIS report sources
  • DICOMweb STOW-RS over HTTPS — for cloud and third-party PACS sources
  • Mutual TLS with SHA-256 client-certificate thumbprint pinning on every channel
  • Optional sealed bearer token as an alternative to mTLS on the DICOMweb channel
  • Off by default — on-prem deployments keep plain DIMSE / MLLP inside the trusted network

De-identification core

The de-identification itself, applied under the HIPAA Expert-Determination method. Identity is replaced deterministically so the same patient stays linkable across studies without any correlation table.

  • Keyed-HMAC pseudonymization of patient and study identifiers (per-site key, never leaves the engine)
  • Interval-preserving per-patient date shift (DICOM Attribute Modification 113107) — exam intervals and chronological order preserved, absolute dates obscured
  • VR-aware tag substitution — replacements always fit the target value representation (no silent blanking)
  • Private / proprietary tag removal (default on)
  • Deterministic study-UID remapping — the de-identified study carries a stable, re-derivable UID
  • Optional operator-chosen sequential labeling (licensed; requires a captured §164.514(c) attestation)

Content gate — free text, pixels, documents

Identifiers hide in more than structured tags. The content gate finds and removes PHI in narrative reports, burned-in image pixels, and encapsulated documents — or quarantines what it cannot safely clean.

  • Free-text SR narrative scrubbing (shared PHI heuristic scrubber)
  • Burned-in pixel OCR redaction — local Tesseract renders the pixels and blacks out detected text (never cloud OCR)
  • Encapsulated-document handling (PDF rasterize + OCR-redact + rebuild, or text scrub)
  • Per-lane policy: clean, quarantine, or reject — operator-configurable per content type
  • Quarantined objects are sealed on the engine for operator review — never silently forwarded

HL7 v2 ORU de-identification

Reports de-identify with the SAME per-patient pseudonym and date-shift offset as the images, so a de-identified report lands on its de-identified study's timeline with no correlation table and in any arrival order.

  • MLLP listener de-identifies inbound ORU^R01 (PID / OBR / OBX / PV1 / ORC)
  • Cross-protocol demographic alignment — HL7 and DICOM derive the pseudonym + date shift from the same MRN + site key
  • OBX narrative scrubbed via the shared free-text PHI scrubber
  • Store-and-forward with fail-closed NACK — the sender holds any message the engine could not fully de-identify and forward
  • Forwards the de-identified report to a downstream HL7 research sink

Coded-concept harvest → research cohorts

Beyond de-identified images, the engine can harvest standards-coded clinical concepts from report-bearing objects into a PHI-free, pseudonym-keyed cohort store — the input to a Clinical Cohort Builder. Opt-in and gated.

  • Extract standards-coded concepts from pre-scrub report objects (never free-text PHI blobs)
  • PHI-free coded-concept batches keyed by the study pseudonym
  • Two-person export control — a licensed add-on (deid_cohort_export) AND a current §164.514(b)(1) Expert-Determination attestation
  • k-anonymity floor enforced on cohort export
  • Fail-open harvest: a harvest or ingest error never affects the de-identified image forward

Operations, durability & administration

A production appliance: durable store-and-forward, sealed configuration and secrets, PHI-free health telemetry, and the shared administration surface used across the fleet.

  • Durable outbound spool — store-then-drain so a downstream outage or restart never loses an instance
  • Quarantine + dead-letter queues with operator review
  • Whole-configuration encryption at rest (SYNTHIMG envelope; reversible secrets sealed, never emitted)
  • Tamper-evident, hash-chained HIPAA audit log with configurable retention
  • Web admin UI with shared sidebar shell + role-based access control
  • PHI-free /api/health telemetry (queue depth, disk, channel posture) for pool + gateway monitoring

Dataflow

The engine is the trust boundary.

Raw PHI arrives from the customer site over an encrypted, mutually-authenticated channel — DICOM-TLS C-STORE, MLLP-over-TLS, or DICOMweb STOW-RS. The engine de-identifies at the boundary and forwards only de-identified output to VNA Research. It occupies the same position whether Synthology hosts it in the cloud or the customer runs it on-prem.

Secure cross-network de-identification ingest A customer site on the left holds the plaintext sources: a PACS or modality that speaks DICOM C-STORE, an EMR or RIS that emits HL7 v2 ORU reports, and a third-party PACS or cloud that speaks DICOMweb. Each pushes to the De-Identification Engine across an untrusted network over an encrypted, mutually-authenticated channel: DICOM-TLS C-STORE, MLLP-over-TLS, or DICOMweb STOW-RS with mutual TLS. Every channel is fail-closed — an unauthenticated or unpinned peer is rejected before any object is read. The De-Identification Engine is the single trust boundary and the only place the re-identification key lives: it pseudonymizes identifiers, interval-shifts dates, scrubs free text and burned-in pixels, and harvests standards-coded concepts. Only de-identified output leaves it, flowing to VNA Research — a Research-Use-Only de-identified archive plus a PHI-free coded-concept cohort store. The engine occupies the same trust-boundary position whether Synthology hosts it in the cloud or the customer runs it on-premises. CUSTOMER SITE SECURE TRANSIT · UNTRUSTED NETWORK DE-IDENTIFICATION ENGINE VNA RESEARCH · RUO PACS / modality DICOM C-STORE · plaintext PHI EMR / RIS HL7 v2 ORU · plaintext PHI 3rd-party PACS / cloud DICOMweb · plaintext PHI DICOM-TLS · C-STORE secure DIMSE · TLS 1.2/1.3 MLLP-over-TLS secure HL7 v2 DICOMweb · STOW-RS HTTPS · multipart Part 10 mTLS · FAIL-CLOSED · PINNED De-Identification Engine The only place the key lives Pseudonymize · interval date-shift Free-text + burned-in pixel scrub Coded-concept harvest Expert-Determination method TRUST BOUNDARY VNA Research De-identified DICOM archive + PHI-free coded cohort store Downstream HL7 sink de-identified ORU ✓ NO PHI LEAVES de-identified Encrypted, mutually-authenticated ingest (raw PHI, in transit only) De-identified output — no PHI, no key Same trust boundary, two deployments: Synthology-hosted (the engine + VNA Research run in Synthology's cloud) · or customer-operated (the customer runs the whole chain on-prem).

Use cases

Five deployment patterns.

Cloud-hosted or on-prem, image or report, external-vendor shielding or research-cohort building — the engine is the single de-identification point for all of them.

Pattern 1

Synthology-hosted research ingest (cloud)

A customer sends studies to a Synthology-hosted De-Identification Engine across the internet. Raw PHI crosses only inside a mutually-authenticated, encrypted channel; nothing re-identifiable is ever stored in the cloud.

Flow

  1. 1 Customer PACS / modality sends over DICOM-TLS C-STORE (or DICOMweb STOW-RS) with a pinned client cert
  2. 2 The engine authenticates the peer, decrypts in memory, and de-identifies at the trust boundary
  3. 3 Identifiers pseudonymized, dates interval-shifted, free text + burned-in pixels scrubbed
  4. 4 De-identified study forwarded to VNA Research; PHI-free coded concepts harvested to the cohort store
  5. 5 No PHI and no re-identification key persist in the cloud — only de-identified output
Pattern 2

On-prem de-identification appliance

The customer runs the whole chain inside their own network. The engine is the internal gate between clinical archives and a research environment — TLS optional, since traffic never leaves the trusted network.

Flow

  1. 1 Clinical PACS / Router forwards studies to the on-prem engine over plain C-STORE (inside the trusted LAN)
  2. 2 The engine de-identifies under the site's Expert-Determination policy
  3. 3 De-identified output lands in an on-prem VNA Research archive
  4. 4 Coded concepts populate the local Clinical Cohort Builder
  5. 5 The re-identification key stays on the customer's engine; researchers see only de-identified data
Pattern 3

AI-vendor / external-reader PHI shielding

Send de-identified studies to an AI vendor or outside reading service that is not in BAA scope. They receive only pseudonymized data; the mapping stays on the engine.

Flow

  1. 1 Source study de-identified with keyed-HMAC pseudonymization + date shift
  2. 2 De-identified study sent to the AI vendor / external reader
  3. 3 The vendor analyzes and returns findings tagged with the same pseudonyms
  4. 4 The re-identification mapping — held only on the engine — links findings back to the right patient
  5. 5 The vendor never received PHI; the clinical workflow still benefits from the findings
Pattern 4

Research report de-identification (HL7)

ORU reports from the EMR / RIS de-identify alongside the images and land on the matching de-identified study's timeline, so a research reporting environment has both the pixels and the narrative — with no PHI.

Flow

  1. 1 EMR / RIS emits ORU^R01 over MLLP-over-TLS
  2. 2 The engine de-identifies PID / OBR / OBX / PV1 / ORC and scrubs the OBX narrative
  3. 3 HL7 and DICOM derive the same pseudonym + date shift from the same MRN + site key
  4. 4 The de-identified report is store-and-forwarded to the research HL7 sink (fail-closed on any failure)
  5. 5 Report and study align on the timeline with no correlation table
Pattern 5

Clinical-cohort building for research

A research program needs cohorts defined by coded clinical concepts, not free text. The engine harvests standards-coded concepts into a PHI-free, pseudonym-keyed store that a Clinical Cohort Builder queries under k-anonymity.

Flow

  1. 1 Harvest lane extracts standards-coded concepts from pre-scrub report objects (opt-in, licensed)
  2. 2 PHI-free coded batches keyed by the study pseudonym are posted to VNA Research
  3. 3 A researcher composes cohort criteria (coded concepts, demographics, intervals)
  4. 4 Export is gated: the deid_cohort_export add-on AND a current §164.514(b)(1) Expert-Determination attestation, both required
  5. 5 Cohorts export under an enforced k-anonymity floor

Integration points

Sources in, de-identified output out.

Sources (senders)

  • Modalities + PACS over DICOM C-STORE (plain on-prem, or DICOM-TLS across a network)
  • EMR / RIS over HL7 v2 ORU (plain MLLP, or MLLP-over-TLS)
  • Third-party PACS / cloud over DICOMweb STOW-RS (mTLS or sealed bearer)
  • XyDromatics Router — can route studies to the engine for de-identification

Destinations (de-identified output)

  • XyDromatics VNA Research — de-identified DICOM archive (Research-Use-Only)
  • VNA Research coded-concept cohort store (PHI-free) — the Clinical Cohort Builder input
  • Downstream HL7 research reporting sink (de-identified ORU)
  • AI vendors / external reading services — pseudonymized studies only

Cross-product orchestration

  • Shared per-site encryption key — consistent pseudonyms across Router, VNA family, and the engine
  • Shared DICOM SCP library — the same C-STORE receive path used fleet-wide (now with optional DICOM-TLS)
  • Unified key-escrow (IKeyEscrow) — every reversible secret sealed under one customer master
  • Synthology QMS — change control on de-identification policy

Identity, audit, observability

  • OIDC / SAML + Active Directory / Azure AD for the web admin UI
  • Hash-chained, tamper-evident HIPAA audit log
  • PHI-free health telemetry (SynthIQ pool + SynthGateway console)
  • Per-peer, per-channel audit of every accepted association

System requirements & sizing

Sized to study volume + content-gate depth.

Pixel-OCR redaction and encapsulated-document handling are the heaviest operations; sizing reflects study volume and how much of the content gate is enabled.

Tier Study volume CPU RAM Storage Typical site
Small (single facility) < 5,000 studies / day 4 vCPU 8 GB 250 GB (spool + quarantine + audit) Imaging center feeding a single research program.
Medium (community hospital) 5,000 – 25,000 studies / day 8 vCPU 16 GB 1 TB Full modality mix with pixel-OCR + HL7 report de-id.
Large (academic / multi-site) 25,000 – 100,000 studies / day 16 vCPU (HA pair) 32 GB per node 2 TB per node Academic center; cohort harvest + multi-source cross-network ingest.
Enterprise (Synthology-hosted / IDN) > 100,000 studies / day 32 vCPU (HA cluster) 64 GB per node 4 TB per node Cloud-hosted multi-tenant research ingest across many customer sites.

Licensing

Three packages.

Core de-identifies and forwards. Secure Ingest & Content Gate adds the cross-network channels, HL7 report de-id, and the pixel / document scrubbing. Cohort Builder adds the coded-concept harvest under two-person Expert-Determination control.

De-ID Core

DICOM de-identification + store-and-forward. The base appliance for turning clinical studies into de-identified output.

  • C-STORE SCP ingest + de-identified C-STORE forward
  • Keyed-HMAC pseudonymization + interval-preserving date shift
  • Private-tag removal + VR-aware substitution
  • Durable spool, quarantine, dead-letter
  • Hash-chained HIPAA audit + web admin UI + RBAC

De-ID + Secure Ingest & Content Gate

Adds cross-network secure ingest (all three channels), HL7 report de-identification, and the content gate (free text + pixels + documents).

  • Everything in Core
  • DICOM-TLS C-STORE + MLLP-over-TLS + DICOMweb STOW-RS (all mTLS)
  • HL7 v2 ORU de-identification with cross-protocol alignment
  • Free-text narrative scrub + burned-in pixel OCR redaction
  • Encapsulated-document handling (PDF / text)

De-ID + Cohort Builder

Everything, plus the coded-concept harvest lane and the Clinical Cohort Builder — under two-person Expert-Determination export control.

  • Everything in Secure Ingest & Content Gate
  • Coded-concept harvest → PHI-free pseudonym-keyed cohort store
  • deid_cohort_export add-on with §164.514(b)(1) attestation gate
  • k-anonymity-floored cohort export
  • Active-active HA pair · premium support eligible

Per-deployment licensing. Multi-product bundles (De-ID Engine + VNA Research + Router) get bundle pricing and shared site-key provisioning. The Synthology-hosted deployment is offered as a managed service — see Managed services.

Documentation

The controlled-document set.

Current GA release v1.0.1.63 · released 2026-08-05

Every De-Identification Engine deployment ships with the documents below, all managed under SynthQMS document control. The DICOM Conformance Statement is publicly available; the rest are shared under mutual NDA.

What’s new·v1.0.1.63

Raw PHI in, research-safe data out — a dedicated de-identification pipeline

  • Secure, fail-closed ingest over mutually-authenticated TLS — DICOM C-STORE, HL7 v2, and DICOMweb STOW-RS — so an unrecognized peer is turned away before any object is read.
  • De-identification designed to support a HIPAA Expert-Determination workflow — identifiers are pseudonymized deterministically, so the same patient stays linkable across studies, and dates are shifted per patient to preserve intervals while obscuring the calendar.
  • A content gate beyond structured tags — scrubbing free-text report narratives, redacting burned-in pixel text with on-box OCR (never the cloud), and handling encapsulated PDFs — quarantining anything it cannot safely clean instead of forwarding it.
  • De-identify HL7 v2 result reports onto the same de-identified study’s timeline, and optionally harvest standards-coded clinical concepts into a PHI-free, pseudonym-keyed cohort store — gated by two-person export control and a k-anonymity floor.
  • Runs Synthology-hosted or on-prem, with the re-identification key kept only on the engine, durable store-and-forward, secrets sealed at rest, and a tamper-evident HIPAA audit log. Ships with the approved EULA and clears a full QA + security-test pass.
Document Title Notes
DOC-2026-444 Hardware & Software Requirements (De-Identification Engine) Sizing, OS support, dependencies
DOC-2026-445 DICOM Conformance Statement (De-Identification Engine) Public — Storage SCP + DICOMweb STOW-RS ingress; de-identified egress SCU
DOC-2026-446 EULA Annex (De-Identification Engine) Annex to the Synthology Master EULA (DOC-2026-063)
DOC-2026-450 Penetration Test Results (De-Identification Engine) Available under NDA
DOC-2026-449 QA Runner Manual (De-Identification Engine) Internal QA + customer-validated test runs
DOC-2026-447 User Guide (De-Identification Engine) Operator + administrator workflows
DOC-2026-448 Installation Guide (De-Identification Engine) Includes MSI deploy + Linux tarball install

Frequently asked

The questions prospects ask.

How is this different from anonymization inside the Router or the VNA?

The Router and the VNA family can pseudonymize on export as part of a routing or archive workflow. The De-Identification Engine is the dedicated, one-way appliance whose whole job is turning raw PHI into a research-safe data set under the HIPAA Expert-Determination method — with the secure cross-network ingress, the content gate (free text + burned-in pixels + documents), HL7 report de-identification, and the coded-concept cohort harvest that a general router does not carry. It is the trust boundary between clinical data and a research environment.

Do you support secure ingest over both DIMSE and DICOMweb?

Yes — both, plus HL7. A source can push over DICOM-TLS C-STORE (secure DIMSE, for C-STORE-only modalities and legacy PACS), MLLP-over-TLS (secure HL7 v2), or DICOMweb STOW-RS over HTTPS. Every channel is mutually authenticated with SHA-256 client-certificate thumbprint pinning and is fail-closed: an unauthenticated or unpinned peer is rejected before any object is read. All of it is off by default, so an on-prem deployment inside a trusted network keeps plain DIMSE / MLLP.

Who holds the re-identification key?

Only the engine. Pseudonymization uses a keyed HMAC with a per-site key that never leaves the engine, sealed at rest under the unified key-escrow master. De-identified output carries no key and cannot be re-identified by any downstream — VNA Research, an AI vendor, or an external reader. In the Synthology-hosted deployment, no PHI and no key persist in the cloud; only de-identified output does.

Does the date shift break longitudinal research?

No — that is the point of it. Each patient gets a stable, per-patient offset (HMAC-derived), applied uniformly to every date. Exam intervals and chronological order are preserved exactly, so longitudinal analysis still works; only the absolute calendar dates are obscured. The engine stamps DICOM Attribute Modification 113107 whenever a date was actually shifted. Ages over 89 are handled per Safe-Harbor expectations.

How do reports stay aligned with their de-identified studies?

Cross-protocol demographic alignment. Both the HL7 and the DICOM paths derive the pseudonym AND the date-shift offset from the same MRN + site key, so a de-identified report lands on its de-identified study's timeline with no correlation table and regardless of which arrives first. The HL7 path is store-and-forward and fail-closed: the sender holds any message the engine could not fully de-identify and forward.

What is the coded-concept cohort harvest, and how is it controlled?

An opt-in lane that extracts standards-coded clinical concepts from report-bearing objects into a PHI-free, pseudonym-keyed store — the input to a Clinical Cohort Builder. It never stores free-text PHI blobs. Turning it on and exporting cohorts is a two-person control: it requires the deid_cohort_export licensed add-on AND a current §164.514(b)(1) Expert-Determination attestation on file, and cohort export is held to an enforced k-anonymity floor.

What is the regulatory classification?

Non-device software under FD&C Act §520(o)(1)(D) — the 21st Century Cures Act carve-out for software that transfers, stores, converts, or displays device data without interpreting it. The engine does not interpret images or reports; it de-identifies and forwards. Its de-identified output is intended for Research Use Only, and the de-identification follows the HIPAA Expert-Determination method (45 CFR §164.514(b)(1)).

Is it cross-platform?

Yes — Windows (MSI installer) and Linux (self-contained tarball), same as the rest of the fleet. The Synthology-hosted deployment runs on Linux; on-prem customers run either.

Build your de-identified research pipeline.

Cross-network ingest to a Synthology-hosted engine, or an on-prem appliance inside your network? Tell us where your PHI starts and where your research data needs to land, and we’ll come back within one business day with a proposed architecture.