Architecture
Route around it, or keep a copy.
Most imaging estates buy availability twice — once in a clustered load balancer that does not understand DICOM, and again in shared storage that has to be built, tuned and kept alive. We do neither, because the right answer is not the same for every product.
There are two mechanisms, and which one applies depends on what a product does with your data.
Products that move data
Route around the failure
The Router and the XyDromatics engines sit in a SynthIQ pool behind a designated standby. They hold nothing that cannot be re-sent, so the answer to a dead node is simply to stop sending to it.
Products that hold data
Keep a copy elsewhere
The archives are the system of record. Routing around one does not help, because what you need is the data itself. Their answer is cross-site replication.
The model
A pool, and a standby that waits.
Backends are grouped into ordered tiers. The primaries are one tier; a designated standby sits in the next. SynthIQ walks the tiers in order and sends traffic to the first tier that has a backend up — so the standby carries nothing at all while any primary is alive.
Modalities point at one address and never learn about any of this. Adding a Router, retiring one, or taking one down to patch it is a change behind SynthIQ, not a change to every modality in the building.
Measured, not asserted
This exact topology is on the benchmarks page
One SynthIQ instance distributing across a two-node Router pool, 20,000 synthetic CT instances per run at 32 concurrent senders, every product at shipped defaults with no tuning. Across four runs every instance was delivered and none failed. A fifth run is excluded rather than averaged in: it shed 350 of 20,007 instances, which we missed initially because we looked at host telemetry instead of the delivery record.
Worth being precise about what that shows: the pooled path carries real load without losing anything. It is not a measurement of switchover time, and we do not publish one.
Detection
Two signals, because one of them lies.
A backend whose web service answers HTTP 200 can still have a stopped DICOM listener. To an HTTP-only health check that node looks perfectly healthy, and every study sent to it fails. So SynthIQ checks both: the backend’s health endpoint, and a real DICOM C-ECHO to its host, port and called AE.
| State | What it means | What SynthIQ does with a bound study |
|---|---|---|
| Unhealthy | A transient HTTP blip. The DICOM listener may well still accept the study. | Honours the existing binding, so the study stays co-located with the rest of itself. |
| Down | C-ECHO confirms the listener is accepting nothing. A forward would fail. | Reroutes to a healthy backend and delivers now, rather than forwarding onto a dead listener. |
Both behaviours are defaults and both are configurable. A site that wants strict co-location can pin a study to its bound backend even while that listener is down, and accept the retries.
Failover and failback
It switches itself, and switches back.
When every primary is down — failed on health and confirmed dead by C-ECHO — the standby tier becomes the active tier. There is no operator action, no cutover procedure and no maintenance window.
And a study is never split by failback.
When the pool recovers, new studies go back to it automatically. But a study that moved to the standby stays bound to the standby until it is done — so its remaining images, and its priors, land in the same place as the rest of it.
The moment of failover itself is the one case that can divide a study: if a node is lost mid-study, the images it had already accepted are on that node and the remainder land on the standby. The affinity binding prevents any further division, and the deployment guide documents the recovery procedure for the split portion — we state this rather than round it away, because it is the difference between a failover story and a fairy tale.
That property is the clinically important one. A failover that returned mid-study would scatter one patient’s images across two archives, and the person who discovers it is a radiologist looking at half a study.
High availability vs business continuity
Same mechanism. The difference is geography.
There is no separate product, no separate licence and no separate configuration for business continuity. You decide where the standby box goes, and that decision alone determines whether you have node resilience or site resilience.
Nor does it stop at two. The tiers are ordered and open-ended, so a three-site estate can run primaries in one data centre, a standby in a second and a further standby in a third, with SynthIQ walking them in order as each is exhausted.
The archives
An archive cannot be routed around.
Everything above protects products that move data. If a Router dies you stop sending to it and nothing is lost, because it was never the system of record.
An archive is the system of record. Routing around it gets you a healthy node with none of your images in it, which is not a recovery. So the archives of record — Repository, Research, Clinical and Clinical Research — are made resilient the only way that actually helps: the data exists in more than one place.
On write
Every instance, every peer
As each instance is archived it is queued for every enabled peer site whose modality filter matches it, and delivered over DICOMweb STOW-RS. A peer can take everything, or just the modalities that site is meant to hold.
As a backstop
A sweep that re-checks
A slow, bounded reconciliation pass walks the archive and compares it against each peer. It is deliberately throttled so it never competes with live ingest — it exists to catch what the write path missed, not to carry the load.
And you can check
A shortfall is visible
The sweep reports how many instances are missing on each peer. That number is the point: replication you cannot measure is replication you are trusting rather than verifying.
Why both paths use one rule
The write path and the reconciliation sweep evaluate the same modality filter with the same predicate. If they disagreed, an instance could be pushed by one and judged out of scope by the other, and the shortfall count would report gaps that were never real — which would make the number worthless precisely when you needed to rely on it.
Transfers retry on failure, and a job left in flight by a crashed process is reclaimed only after a delay long enough that a merely-slow transfer is never mistaken for a dead one and sent twice.
Worth being plain about: the archives are not members of a SynthIQ pool. Replication is a data-protection mechanism, not a request-routing one — it means a second site has your images, not that a failed archive is transparently substituted mid-association. If your requirement is continuous archive availability rather than durability, that is a design conversation, and we would rather have it than let you infer a capability from a page.
The obvious question
So what happens if SynthIQ goes down?
Fair question, and the honest answer is that you should run it as a pair. Two or more SynthIQ nodes share an external affinity store — PostgreSQL or SQL Server rather than host-local storage — so the record of which study went where lives outside any one node.
Because affinity is external, losing a node loses neither the affinity nor the in-flight work: the surviving node already knows where every study was going.
You can confirm this rather than trust it. SynthIQ’s health endpoint reports whether its affinity store is external, so an HA-configured deployment is verifiable at a glance instead of being a matter of documentation.
Where it stops
We will not tell you there is no single point of failure.
There always is one. Every layer you harden moves it somewhere else rather than removing it, and any vendor who claims otherwise is selling you the next layer. The useful question is not “is it gone?” but “where is it now, and can we live with it there?”
| If this is your weak point | You harden it like this | And the weak point moves to |
|---|---|---|
| A single archive | Replicate cross-site to a peer archive. | Your images exist in two places. The link between the sites, and the storage at the peer site, are now what has to stay healthy. |
| A single Router | Put two or more Routers in a SynthIQ pool. | The pool survives losing a node. SynthIQ is now the thing that has to stay up. |
| A single SynthIQ | Run SynthIQ as a pair over a shared external affinity store. | Either node can answer for any study. The affinity store is now the thing that has to stay up. |
| The affinity store | Run PostgreSQL or SQL Server in whatever HA configuration your DBA team already operates. | The database cluster survives a node loss. Its quorum and its network are now the thing that has to stay up. |
| The site | Put the standby tier in a second data centre. | You survive losing the primary site. Your WAN between the sites is now the thing that has to stay up. |
Most sites stop at the second or third rung, and that is a reasonable place to stop. We would rather show you the ladder and help you pick a rung than pretend the top of it is a place you can reach.
What we will not do is sell you a clustered load balancer in front of the load balancer. That is the architecture SynthIQ exists to replace.
What you do not have to build
No shared filesystem. No cluster. No quorum.
In a pool, each Router or engine owns its own queues and its own storage. There is nothing to fail over by hand, no cluster membership to manage, and no split-brain condition to reason about, because the pool members do not coordinate with each other at all — SynthIQ decides, in front of them, where each study goes.
That is why availability here is a deployment choice rather than a project. The work is standing up one more node and telling SynthIQ which tier it belongs to.
Questions
The ones that come up.
Is high availability an extra licence or a higher tier?
No. It is included with every licence. There is no HA edition, no HA feature flag, and no upgrade path to buy. What varies is the engagement: designing the placement, the network path and the failover drill is professional-services work, and operating it for you is the managed-service tier.
How is this different from putting a load balancer in front of the pool?
A general-purpose load balancer distributes connections. It has no idea that forty images belong to one study, so it can scatter them across backends and leave a study half-delivered in two places. SynthIQ routes on Study Instance UID and keeps a study together on one backend, which is the property that matters clinically. It also verifies the DICOM path rather than just an HTTP port.
Does a floating IP give me the same thing?
No. A floating IP is address ownership at the network layer — it moves an address between hosts. It does not know which backend a study was already going to, cannot tell a healthy web service from a stopped DICOM listener, and has no concept of study affinity. It is a useful complement to SynthIQ, not a substitute for it.
What happens to a study that is mid-transfer when a backend dies?
It depends on how the backend died, and the distinction is deliberate. A transient HTTP blip leaves the DICOM listener able to accept the study, so the existing binding is honoured and the study stays co-located. A listener confirmed down by C-ECHO would fail the forward, so the study is rerouted to a healthy backend and delivered now rather than dying on a dead listener.
Can I have more than one standby tier?
Yes. Tiers are ordered and open-ended, so a three-site estate can run primaries in one data centre, a standby in a second and a further standby in a third. SynthIQ walks the tiers in order and uses the first one with a backend that is up.
Why are the archives not in a SynthIQ pool?
Because routing around an archive does not solve the problem. A pool works when the members are interchangeable, and they are interchangeable only because they hold nothing you cannot re-send. An archive is the system of record, so a healthy substitute with none of your images in it is not a recovery. Archives are protected by cross-site replication instead, which puts the data itself in a second place. If you need continuous archive availability rather than durability, that is a design conversation to have with us rather than something to assume from a diagram.
Do I need shared storage between the Routers?
No, and this is the part people expect to be harder than it is. Each Router owns its own queues and its own storage. There is no shared filesystem to build, no cluster membership, and no quorum to lose. What makes the site resilient is SynthIQ in front of it.
Work out where your weak point is.
Bring us your topology and we will tell you honestly which rung of the ladder you are on, what the next one costs, and whether it is worth climbing.