Whitepaper · Capacity planning for radiology IT

Right-sizing DICOM routing for mammography.

A measured approach to capacity planning for XyDromatics Router and SynthIQ. We put the software on a dedicated performance rig and measured exactly what drives capacity — so you buy the right box, not the most expensive one. The headline, corrected by re-measurement on real image objects: routine-imaging delivery is storage-bound — the spool volume’s write path sets the ceiling and added cores do not move it — while 3D tomosynthesis memory use is a configured ceiling (~75% of host RAM) working as designed.

Download PDF Synthology Healthcare Solutions Group · Published 2026-06-03 · ~12 min read

Executive summary

Sizing a DICOM router is usually guesswork: a vendor quotes "enterprise-grade hardware," a hospital over-buys storage it never uses, and nobody can say what actually happens when four mammography units start pushing 3D tomosynthesis studies at once.

We took a different approach. On a dedicated performance rig — an isolated DICOM generator driving a system-under-test across three host sizes — we measured exactly what governs XyDromatics Router capacity. The findings are clear, and a few are counter-intuitive:

  • Routine imaging is storage-bound, and cores do not move it. Measured on real 512 KB CT instances (shipping build, spool on ext4, six replicates): ~155 instances/sec at 2, 4 and 8 vCPU alike — 4→8 vCPU scaling of 1.00×. On the same host, moving that spool to a RAM filesystem took ingest to 849 instances/sec — 6.50× — so the spool volume’s write path is the ceiling. The ~700/~1,500 figures previously headlined here are real retained measurements — of 1 KB test objects, roughly 500× smaller than the CT instances this page discusses; they were mislabeled as study-instances/sec and are re-scoped in §4.
  • Ingest is not delivery — size on delivery. A router answers the modality as soon as it has spooled the object; forwarding runs behind that queue. On the two 8 vCPU hosts where we measured delivery directly, the same workload delivered at 159 and 196 instances/sec against ingest of 511 and 751 — 3.2× and 3.8× lower (§4.1). On a disk-backed spool the two clocks converge instead (measured 1.09× on ext4, §4.1) — so an ingest rate overstates capacity roughly threefold on a RAM-backed spool, and barely at all when real storage gates both legs.
  • Provision storage write bandwidth, not IOPS. The June finding that quadrupling IOPS changed nothing was an ingest-side observation and survives — but it was read as “disk speed is irrelevant,” and that reading is wrong. For delivery, the spool volume’s sustained small-file fsync bandwidth is precisely the limit (§4): a datasheet IOPS number does not predict it, a measured fsync throughput does.
  • 3D mammography’s memory use scales with the host, by design. The 75% heap ceiling adapts: the same 16-volume workload peaked at ~11 GB under a 16 GB cap, 21–24 GB under 32 GB, and 41–43 GB uncapped on 62 GB — accepted in full at every tier, with latency (not refusals) as the cost of a smaller box in our re-measured runs (§5).
  • High memory use under load is by design, not distress — now confirmed by direct measurement. The router’s runtime is configured with a hard heap ceiling at 75% of host RAM; the heap grows toward that ceiling and recovers afterward. Removing the cap on a 62 GB host nearly doubled peak use (21.7→41.2 GB at the same load), proving the ceiling was what bounded it (§5–§6).
  • Capacity grows by adding nodes, not bigger boxes — measurably, though not linearly. A SynthIQ pool load-balances across routers and places large studies on nodes with memory to spare; two nodes delivered 432 instances/sec against 350 for one, and 0.65 of the two measured separately (§7).

The result is a sizing matrix you can buy against with confidence — with the rate expectations stated as delivery rather than ingest — and a scale-out trigger that tells you exactly when to add the next node.

§1 — Why router sizing is misunderstood

Two myths drive most over-spending.

Myth 1: "The datasheet IOPS number tells you what you need."

An earlier version of this section said storage speed doesn’t matter at all — the router barely used a modest disk’s IOPS budget. That was an ingest-side observation, and for delivery it is the wrong lesson. Re-measured on real 512 KB CT objects, the router’s rate tracks the spool medium and nothing else — swapping that medium moves it 6.5× while quadrupling the cores moves it under 1%, so storage is exactly the limit. What stays true is that the datasheet number is the wrong one: random-IOPS ratings do not predict sustained fsync-per-file throughput, which is what a DICOM spool actually does. Measure the volume with a small-file fsync test, or ask us to.

Myth 2: "Capacity is one number."

A router's ceiling depends entirely on the shape of the traffic. A stream of small CT and CR images stresses completely different resources than a handful of 561 MB tomosynthesis volumes. Quoting "X studies/hour" without specifying the modality mix is meaningless. Real sizing needs a model, not a number.

§2 — Methodology

Credibility here comes from rigor, so the method matters.

  • Dedicated rig, isolated roles. One VM generated DICOM load; a separate VM ran the router under test; they communicated over a private network. Nothing else competed for resources.
  • Three host sizes. The system-under-test was measured at 4 vCPU / 16 GB, 4 vCPU / 32 GB, and 8 vCPU / 64 GB — so we could see how the limits move with the box, not just read one data point.
  • Realistic send patterns. Routine load used study-mode — many images per DICOM association, the way real modalities transmit — not the pessimistic one-association-per-image worst case.
  • Dual-host attribution. Every high-load run sampled CPU on both the generator and the router. This is the discipline that makes the conclusions trustworthy: when 3D mammography began failing at high concurrency, the naïve reading was "the router is overloaded." The data confirmed it on a corrected reading — and showed the generator was idle while the router was CPU-saturated.
  • What the throughput clock measures. Our load harnesses compute a rate as sender-side C-STORE successes divided by the elapsed send loop. A router returns success as soon as it has handed the object to its processing queue; the forward to the destination runs asynchronously behind that queue. The stopwatch therefore stops well before delivery, and every rate produced this way is an ingest acceptance rate. Delivery has to be counted on the receiving side against its own clock — §4.1.

§3 — The central finding: a bottleneck that moves

The headline result is still a two-regime model — but re-measurement moved one of the levers. For everyday imaging the lever is the spool volume’s write bandwidth: measured on real CT objects, delivery was identical at 2, 4 and 8 vCPU, while swapping the spool medium moved it 6.5× (§4). For 3D mammography the working set is governed by a configured memory ceiling at 75% of host RAM (§5). A tomosynthesis-heavy site sizes RAM; every site sizes the storage write path; cores follow the transform load (§4.1’s realistic mixed-modality workload is where CPU genuinely binds), not the raw routine-imaging rate.

What actually limits 3D tomosynthesis throughput Two measured findings for 3D tomosynthesis on XyDromatics Router, each volume about 561 megabytes. First: memory is not a wall. Peak process memory settles at roughly 67 to 70 percent of whatever memory limit the process is given — 10.8 gigabytes under a 16 gigabyte limit, 21.4 under 32, and 43 on an unconstrained 62 gigabyte host — because the .NET heap honours a configured hard limit of 75 percent. Every volume was accepted at every limit and every concurrency tested; nothing was shed and nothing ran out of memory. Second: latency is the constraint that actually moves. It grows close to linearly with concurrency, from 24 seconds per volume at 4 concurrent to 102 seconds at 16 concurrent, while acceptance stays at 100 percent throughout. Sizing for tomosynthesis is therefore a latency decision, not a memory-headroom decision. What actually limits 3D tomosynthesis throughput Memory is a ceiling, not a wall peak process memory vs the limit it was given 10.8 GB 16 GB limit 68% 21.4 GB 32 GB limit 67% 43.2 GB 62 GB host 70% Zero volumes shed at any limit, at any concurrency tested. Latency is what actually moves median seconds per volume — acceptance stayed 100% 24 s 4 8/8 49 s 8 16/16 75 s 12 24/24 102 s 16 32/32 Concurrent volumes (below) · accepted / offered in green. Size tomosynthesis for the latency you can accept — not for memory headroom. DBT = 3D tomosynthesis, ~561 MB per volume, sent as one object. Concurrency = mammo units sending at once. Measured on one 8 vCPU / 62 GB host, Router 1.2.3.368, spool on ext4. Two replicates per point. Smaller figures were produced by constraining memory on that host, not by measuring smaller hosts.

Measured across memory tiers (re-run 2026-08-25, cold heap per point): the plateau at 40–43 GB on a 62 GB-usable host is the 75% heap ceiling engaging, and the same ceiling scales down with the host — a 32 GB cap truncates the same workload at 21–24 GB and a 16 GB cap at ~11 GB, in every case with all volumes accepted. Smaller hosts run the same load in less memory at the cost of latency, not of refusals — see §5.2.

§4 — Routine imaging: the storage write path sets the ceiling

These are ingest-acceptance rates — what the front door accepts and spools, which is the clock this section is about. What the router then delivers downstream is a separate measurement, in §4.1 below. Re-measured 2026-08-25 on real 512 KB CT instances — shipping build, spool on ext4, CPU affinity verified on every tier, six replicates (RI-2026-644, retained evidence bundle):

vCPU Ingest acceptance (512 KB CT) Accepted
2155.0–155.6 /sec20,000 / 20,000 × 2
4154.6–155.1 /sec20,000 / 20,000 × 2
8154.2–154.6 /sec20,000 / 20,000 × 2

Doubling cores twice changed nothing — 4→8 vCPU scaling is 1.00×. What does move the number is the storage underneath it. On the same host and the same build, moving the spool from ext4 to a RAM-backed filesystem took ingest from 130.6 to 849.1 instances/sec — 6.50×. Across every host and medium measured the sustained rate spans 67 MB/s to 450 MB/s, entirely on what the spool writes to. Cores cannot matter when the disk is the wall.

A withdrawn figure, then measured. An earlier version of this section stated that the volume delivered 80 MB/s for fsync-per-file writes against 169 MB/s raw sequential, and that the router therefore ran at “97–98% of the storage ceiling”. Those figures existed nowhere in the evidence corpus as a measurement — no conditions record, no command line, no replicate count, no output file — so they were withdrawn on that basis. The device has since been measured properly, across three replicates.

Device test Durable writes/sec ms per write MB/sec
raw sequential + fdatasync — — 168.7
520 KiB, fsync per file 168–178 5.6–5.9 85–91
1 MiB, fsync per file 143–153 6.6–7.0 143–153

The raw-sequential figure came back at 168.7 MB/s against a withdrawn claim of 169 — it was right, and had simply never been retained. The fsync figure was close too, and should still not have been quoted, because it is not a device constant: the cost is per file rather than per byte, at 5.6–7.0 ms per durable write at either size, so megabytes per second scales with object size and one number cannot describe it.

The ratio was unsound for a separate reason, which measurement now quantifies. It divided bytes on the wire by bytes on the disk; the router’s write amplification is 2.22× — 1,184 KB reaches the block device per 522 KiB object once database, WAL and index writes are counted — so the two quantities were never comparable. In the unit where both sides are durable writes it is a real number: the router sustains about 159 object-writes per second where the device manages 168–178, or 89–94%, while writing 2.22× the object bytes per write. The original 97–98% was in the right neighbourhood, for the wrong reason and in the wrong unit.

What happened to ~700 and ~1,500 instances/sec? An earlier version of this page reported those figures and later said their raw data had not been retained. Both statements were wrong in instructive ways. The data was retained — in a June results tree nobody re-checked — and the runs are real: 36,000 of 36,000 objects at 1,500.7/sec. But the retained run records pixel_bytes: 1024 — one-kilobyte test objects, roughly 500× smaller than the CT instances the rest of this page discusses, published under a label (“study-instances/sec”) that implied otherwise. At one kilobyte an object the spool write path is nowhere near saturated, so the association-cap and CPU effects the June notes describe are what remains visible — which is how a cores-scale-linearly narrative could be honestly measured and still be wrong for real imaging. The August figures above, on real image objects rather than kilobyte placeholders, are the ones to size from.

One June observation does survive at every object size: more connections aren’t faster. Throughput peaks at moderate concurrency; a client that opens far more simultaneous associations than the configured cap just queues the excess. A client-tuning consideration, not a hardware limit.

4.1 Ingest is not delivery

The figures above are what the front door accepts. They are not what the router delivers. A router answers the modality with success as soon as it has handed the object to its processing queue; the forward to the destination runs asynchronously behind that queue, and the harness clock stops at the last sender-side response. Everything the router has accepted but not yet forwarded is still sitting in that queue when the number is taken.

In August 2026 we measured delivery directly for the first time, counting on the receiving side: one instance per object the router forwarded and the sink acknowledged, clocked from first arrival to last delivery, on a clean database with the queues verified empty at settle. Three runs per host, zero failures, the published workload unchanged (15-second ramp, 500-instance warm-up, 32 concurrent senders).

Host (8 vCPU) Ingest acceptance Delivery Ingest ÷ delivery
EPYC-Rome, 1996 MHz511/sec (488–547)159/sec (154–164)3.21×
EPYC-Milan, 3250 MHz751/sec (724–767)196/sec (195–197)3.83×

Two different ratios live around this table; they are not interchangeable. The last column is the product's own ingest-to-delivery relationship — the re-measured ingest divided by the delivery beside it, 3.21× and 3.83×. The separate question of how far our previously published figures overstated delivery uses those published figures as the numerator instead: 527 against 159 is 3.31×, and 817 against 196 is 4.17×. The first pair describes the product; the second pair describes the size of the error a reader of the old page was handed. We state which is which wherever either appears.

Put the other way round: at the moment the ingest clock stops, 8.0% of the work has been delivered on the Rome host and 2.0% on the Milan host. The rest is queued.

The forwarding leg does not scale with the processor the way the front door does. Those two hosts have the same core count and differ in processor generation and clock. At the front door they are far apart — 751/sec against 511 — and on delivery only 1.23× apart, 196 against 159. Buying a faster processor moves the front door far more than it moves the forward — the opposite of what an ingest-only measurement would lead you to expect. (An earlier revision of this paragraph put the ingest gap at 1.63×, a multiple carried over from a working note that does not divide out of the measured means. It has been removed rather than replaced with a recomputed one; the means themselves are quoted instead.)

For reference, a deliberate control — the same workload on an 8 vCPU Azure host with the spool on tmpfs, which is not a customer configuration — delivered 336/sec at association caps of 40/40/40 (four runs, 329–351) and 321/sec at caps of 32/10/3 (three runs, 316–329), with the forwarding leg measured on its own at 524 and 480/sec. Ingest on that host ran ~724–755/sec, so even with the spool in RAM the delivered fraction at the moment the ingest clock stops was 19.4%.

These delivery runs were made on the cloud rigs of DOC-2026-490. How the §4 ladder relates to them is now measured directly, on the ext4 host itself: ingest 154.5/sec against delivery 141.2/sec — a ratio of 1.09× (two replicates, both reconciling exactly at 20,000 instances, zero refusals). With the spool in RAM the same ratio is 3.21× and 3.83× on the two cloud rigs. When storage binds, the two clocks converge — the same device gates both legs — so a site whose spool is on disk can size from either figure. A site whose spool is in RAM cannot: there, an ingest number overstates delivery roughly threefold. As a cross-cloud sanity check, the two delivery figures on comparable 8 vCPU silicon — 141.2 on Azure ext4 and 159 on the RAM-spool Rome host — agree within 11%.

§5 — 3D mammography: where the bottleneck moves

A single 3D tomosynthesis volume is ~561 MB to over 1 GB — one such study is larger than thousands of routine images combined. This is the sizing-critical case.

5.1 The per-volume memory model — and what it actually predicts

Re-measurement verdict (2026-08-25): the memory model holds — and the apparent 3× contradiction between it and the table below was the measurement resolving two different things. An earlier revision withdrew the per-volume arithmetic because it under-predicts the table’s 40–43 GB cells. The controlled re-run explains both: the router’s runtime carries a hard heap ceiling at 75% of host RAM, and under sustained 3D load the heap grows toward that ceiling regardless of the (much smaller) live working set. The table measures the ceiling; the per-volume arithmetic approximates the working set. Proof the ceiling is what binds: capping the same host at 32 GB truncated the same workload at 21.7 GB peak; removing the cap took the identical load to 41.2 GB — with every volume accepted both times. Size hosts from the table (it is what the process will actually occupy); use the per-volume arithmetic only to reason about relative workload weight, never as a provisioning floor.

5.2 On a 64 GB host the ceiling stops binding — and latency becomes the cost of concurrency

This was the key measured result:

Concurrent DBT volumes Accepted (both replicates) Router CPU (8 vCPU) Latency, median Peak memory
48 / 8 and 8 / 875% / 70%24.1 / 24.4 s19.9 / 20.3 GB
816 / 16 and 16 / 16100% / 100%49.1 / 49.3 s36.7 / 36.5 GB
1224 / 24 and 24 / 24100% / 100%75.5 / 73.4 s43.1 / 43.3 GB
1632 / 32 and 32 / 32100% / 100%102.2 / 101.0 s43.0 / 43.3 GB

Conditions. Two replicates per point on one 8 vCPU / 62 GB host, Router 1.2.3.368, spool on ext4, genuine 561 MB monolithic volumes, cold heap before every point, process RSS sampled directly, CPU sampled on both the generator and the router. Generator CPU stayed at 26–28%, so these figures measure the product and not the harness.
Every volume was accepted at every concurrency, in both replicates. Nothing was shed and nothing ran out of memory — including under a deliberately tight 16 GB memory limit, which accepted the same 32 volumes at a 10.8 GB plateau.
The replicates agreed everywhere. At 12 concurrent their 95th-percentile latencies differ by 17 milliseconds, so these are reproducible figures rather than a single favourable run.

Two things stand out, both now settled by the re-measurement. First, the 40–43 GB plateau is the configured 75% heap ceiling being honoured, not headroom — the same workload truncates at ~11 GB under a 16 GB cap and 21–24 GB under a 32 GB cap, accepted in full every time. Second, past the comfortable band the cost of added concurrency is latency: at 12 and 16 concurrent volumes the run takes visibly longer per volume while everything is still accepted. Nothing was refused at any concurrency measured.

Bottom line: on 8 vCPU / 64 GB, the comfortable band is ~8 concurrent volumes — roughly one per mammography unit sending at once — with rising latency, not refusals, beyond it in our re-measured runs.

§6 — "Why is the router using most of my RAM?"

Under 3D load, the router's memory footprint rises toward ~75% of host RAM — even at light concurrency. This is intentional, and it is now directly measured: the runtime is configured with a hard heap ceiling at 75% of host RAM, and the server garbage collector trades memory for throughput by letting the heap grow to that ceiling before reclaiming. The footprint reflects allocation high-water plus not-yet-collected scratch memory, not the live working set (which is far smaller — capping the same host at half the RAM ran the same workload in half the memory, with nothing refused). It recovers fully after a burst.

The practical guidance: high memory use is healthy as long as the node isn't shedding, and you should not co-locate other memory-hungry services on a mammography router — it will use the RAM you provision for it, by design.

§7 — Growing capacity: scale out, don't scale up

A single router scales up to its limits — cores for routine volume, memory-then-cores for 3D mammography. Past that, the answer is not a bigger box. It is a SynthIQ pool.

SynthIQ load-balances DICOM traffic across multiple routers and routes weight-aware: it predicts a study's size from the modality and SOP class of its very first image, and steers large tomosynthesis studies to backends that have the memory headroom to hold them — keeping the pool balanced and protecting every node from the out-of-memory failure mode. Capacity grows by adding nodes, with no forklift upgrade — and we can now say by how much, because we measured it.

Two routers on identical 8 vCPU silicon, each forwarding through its own outbound leg to its own destination, three runs per condition, every router-class host including both destinations reset to a clean database first, every run settling at exactly 20,000 delivered with zero failures:

Measured Node A alone Node B alone Sum of the two Through the pool Pool ÷ sum
Ingest acceptance858/sec864/sec1,722/sec770/sec0.45
Delivery350/sec317/sec667/sec432/sec0.65

Both rows come from the same runs, and the ingest row reproduces the 0.45 pooling ratio we had already published — which is what makes the delivery row trustworthy rather than merely new. They do not say the same thing. Pooling costs about a third of delivery throughput, not more than half: the 0.45 is an ingest ratio, and using it as "the cost of pooling" understates the pool. And pooled delivery is 1.24× a single node — 432 against 350 — so a second node adds delivered throughput as well as availability. The split across the two nodes was even to within 1%, so balancing is not the constraint. Scale-out is therefore predictable but not linear: plan on the measured ratio, not on multiplying a single node's rate by the node count. Two nodes is what was measured; a third is not.

The reason the two rows differ is structural. Ingest is the front door, and the balancer hop in front of it costs real acceptance capacity. Delivery is the pipeline behind that door, and the pipeline is where a second node genuinely helps — each router forwards through its own outbound leg. On the ingest framing pooling looked like a tax paid for availability; on the delivery framing it adds throughput too. Measuring only the front door hid the benefit the pool exists to provide.

Add a pool node when, sustained: routine throughput approaches the node's ceiling; or concurrent 3D studies approach the node's comfortable band (CPU climbing past ~1.5× core count, latencies stretching); or the node begins to shed gracefully under DBT load. When two of these recur, scale out.

The ceiling to watch is the delivered rate, not the accepted one. A node still accepting at its published ingest rate is not evidence of headroom: acceptance is coupled to the queue behind it, but it is not a measure of the forward, and the two differ by roughly threefold (§4.1). Watch the processing-queue depth alongside the delivered instance count over the same interval — if the queue trends up across a shift, the forwarding leg is the constraint and another node is the answer.

§8 — Recommended configurations

Site profile Routine load 3D mammography Per router
Smalllow–moderatenone4 vCPU / 16 GB / SSD
Mediummoderate–high2D FFDM + occasional DBT (≤4 concurrent)8 vCPU / 32 GB
Large (mammography)high4–8 units, 2D + 3D DBT8 vCPU / 64 GB
  • 32 GB is the floor for 561 MB-class tomosynthesis — a deliberately conservative figure: re-measured, a 32 GB ceiling ran 16 concurrent volumes with nothing refused, at higher latency (§5.2).
  • 64 GB / 8 vCPU is the recommended mammography spec — clearing the memory regime and providing the cores that then become the limit.
  • Size the spool volume by measured write bandwidth, not IOPS. Routine-imaging throughput tracks the spool medium and nothing else: on one host, swapping ext4 for a RAM filesystem moved it 6.5× while quadrupling the cores moved it under 1%. A volume’s IOPS rating does not predict fsync-per-file bandwidth; measure your own volume, or ask us to.
  • Beyond a single node's band, add a SynthIQ pool node.

Size against delivery, not ingest. This table says which resource runs out first — cores for routine imaging, memory-then-cores for 3D mammography. It does not carry a studies-per-hour promise, and any such promise derived from the ingest rates in §4 would over-state a site's real capacity by roughly threefold. The delivery figures in §4.1 — 159 and 196 instances/sec on the two 8 vCPU hosts we measured — are the ones to plan against, and for a two-node pool the measured delivered figure is in §7. We have not measured delivery on the 4 vCPU / 16 GB or 8 vCPU / 64 GB configurations above and will not extrapolate it; ask for a sizing run against your own modality mix instead.

Conclusion

Capacity planning for DICOM routing does not have to be guesswork — but it does have to be re-measured when the evidence improves, and ours did. The corrected statement of what drives capacity: the spool volume’s write bandwidth for routine imaging, a configured 75%-of-RAM ceiling for 3D mammography, and cores for transform-heavy mixed workloads. Customers can buy the right box — not the most expensive one — and know in advance exactly when to add the next node. It also means being plain about which number is which: a router accepts at roughly three times the rate it delivers at, and only the delivered rate is a site’s capacity. That is the difference between a specification and a measurement.

Planning a mammography deployment?

We'll size your Router + SynthIQ deployment against your real modality mix and growth plan — no over-spend, no guesswork.