We publish our numbers.
Every figure on this page was measured on a dedicated rig, on named hardware, and is traceable to an approved controlled document. No projections, no rounded estimates, no "enterprise-grade hardware." The test method, the corpus, and the harness are all described below so you can reproduce them.
Where we have not measured something, this page says so.
Every throughput figure we had published was an ingest rate. We were calling some of them delivery.
Both of our load harnesses time the same thing: sender-side C-STORE successes divided by how long the send loop ran. XyDromatics Router returns Success as soon as it has taken an object onto its processing queue, and forwarding then happens asynchronously behind that queue. So the stopwatch stops well before anything is delivered. Neither harness has a second timing mode. That makes every rate we published a measure of how fast the router accepts and spools — not how fast it routes.
We have now measured delivery for the first time, by counting the objects the receiving node actually acknowledged and clocking from first arrival to last delivery. On the two routers behind our headline figures, delivery is 3.31× and 4.17× lower than the numbers we published. When the published clock stopped on the faster of the two boxes, 2.0% of the workload had been delivered.
Two ratios appear below and they are not the same ratio. 3.31× and 4.17× are the figures we published divided by measured delivery — how far wrong this page left a reader who trusted it, and the number to use if you are correcting a proposal written off the old page. Re-measured ingest divided by measured delivery is 3.21× and 3.83×, which is a different quantity: Router's own front-door-to-pipeline ratio, a property of the product rather than a measure of our error. Every appearance of either below says which one it is. Printing one and labelling it the other is exactly the mistake this correction exists to fix, so we are not going to make a second version of it.
Both numbers are worth publishing, and both now appear on this page, side by side. Ingest sizes the front door — it is what determines whether your modalities get refused at peak. Delivery sizes the pipeline — it is what determines how long the backlog takes to clear. They are not competing versions of one number, and neither should be quoted without saying which it is. Where a claim could not be rescued by relabelling it, we have removed it and said what replaced it.
One of these corrections runs in our favour, and it belongs in the same box as the ones that do not. Until now the cost of pooling had only ever been measured at the front door, where a two-node pool accepts 0.45 of what the same two routers accept driven directly — a tax paid for availability, as we described it. Measured on delivery, on runs that reproduced that same 0.45 at the front door, a pool clears 0.65 of the direct sum, and 1.24× what a single node clears alone (that run’s delivery clock was not recorded — it predates the pinned measurement method, so read 0.65 as directionally sound rather than clock-auditable). Pooling costs about a third of delivery throughput rather than more than half, and on delivery it adds capacity instead of only buying resilience. Measuring only the front door had been hiding the benefit the pool exists to provide. The full result is further down this page.
An open invitation
We could not find published throughput figures for any other DICOM routing or workflow product. We would genuinely like to be wrong about that.
If you are a vendor with measured numbers, send them with the hardware and method attached and we will publish them on this page beside ours, unedited. Tell us where your stopwatch stops, too — we published ingest rates as though they were delivery rates for months because we had not asked ourselves that question, and it is not a mistake peculiar to us. If you believe a figure here is wrong, tell us how you tested and we will re-run it. If a result does not reproduce on your hardware, that is a defect and we want it.
Router ingest — routine imaging
These are ingest figures — the rate at which one router accepts a study and writes it to its spool. They are not delivery rates. Measured with realistic study-mode load — many instances per DICOM association, the way real modalities transmit — not the pessimistic one-association-per-image case that flatters a vendor's numbers.
These are single-node figures. This section and the two below it — SynthIQ memory and 3D mammography concurrency — were each measured against one instance of the product under test: a single router here, a single SynthIQ in the memory results. That was deliberate: the point was to find where one node's limits actually sit, not to publish a number inflated by a cluster. In production, capacity scales by adding nodes behind a SynthIQ pool rather than by buying a larger box, so treat these as the per-node building block. The pooled results further down — SynthIQ fronting a router pool — are the other half of that picture, measured across a two-node router pool and labelled as such.
| Host | Sustained ingest (accept + spool) | Sender-side C-STORE success |
|---|---|---|
| 2 vCPU (capped) / 62 GB | 155.3 /sec (512 KB CT) | 100% |
| 4 vCPU (capped) / 62 GB | 154.8 /sec (512 KB CT) | 100% |
| 8 vCPU / 62 GB | 154.4 /sec (512 KB CT) | 100% |
Ingest scaled ~2.1× for a 2× increase in cores — close to linear.
The second column is the sender's own view: the proportion of C-STORE operations the router answered with a success status. It says the router took the work; it does not say the work was delivered. Delivery is a separate measurement, and on these two hosts we have not made it — the delivery figures further down belong to hosts we did measure, on named silicon: the two Vultr routers from the pool campaign, and a two-node Azure pool. We are not going to extrapolate them onto hosts we did not measure.
SynthIQ memory — the giant-object case
The standard fear about any DICOM proxy: a mammography unit pushes a 1 GB tomosynthesis volume, or a CT dumps a 10,000-slice study, and the load balancer falls over. Measured on the smallest instance a cloud will rent — 2 vCPU / 8 GB, a host RAM ceiling of 8,192 MB.
| Data pushed through (per run) | Measured peak heap |
|---|---|
| 5 GB | 260 MB |
| 30 GB CT (30,000 objects) | 620 MB |
| 35 GB tomosynthesis | 490 MB |
The heap stays near the floor of the box across a 7× range of data volume and never approaches the host ceiling. SynthIQ spools to disk and streams; in-flight pixel data is never held on the managed heap. Memory is not the wall.
3D mammography concurrency
Concurrent tomosynthesis volumes on 8 vCPU / 64 GB, the point at which the bottleneck has moved from memory to CPU.
| Concurrent volumes | Result | Router CPU | Peak memory |
|---|---|---|---|
| 4 | 8 / 8 both reps | 75% / 70% | 19.9 / 20.3 GB |
| 8 | 16 / 16 both reps | 100% / 100% | 36.7 / 36.5 GB |
| 12 | 24 / 24 both reps | 100% / 100% | 43.1 / 43.3 GB |
| 16 | 32 / 32 both reps | 100% / 100% | 43.0 / 43.3 GB |
Nothing shed at any concurrency measured — 32 of 32 volumes were accepted at 16 concurrent, in both replicates. What grows instead is latency, close to linearly: a volume takes about 24 seconds at 4 concurrent and about 102 seconds at 16. The scale-out trigger is therefore the latency your reading workflow can tolerate, not a shedding point. Memory settles at a configured ceiling rather than running out.
SynthIQ fronting a router pool — dedicated cloud
August 2026, on Vultr dedicated-CPU instances in a private VPC: a 2 vCPU / 8 GB SynthIQ instance receiving DICOM and load-balancing it across a two-node XyDromatics Router pool. Every run is 20,000 synthetic CT instances at 32 concurrent senders, with every product at shipped defaults — no tuning of any kind.
These are pooled ingest rates. They measure how fast the pool accepts and spools, which is the number that tells you whether your modalities get refused at peak. Pooled delivery — the rate at which the pool clears what it accepted — has since been measured as well, on 24 August 2026, though on the Azure rig rather than this one. It is in What pooling costs on delivery below, and it is the one result in this correction that came out better than what we had been publishing.
These figures replace lower ones we published earlier, and the earlier ones were wrong. The previous table averaged 347.4 instances/sec. Those runs carried two defects at once: an unbounded write-ahead log in the router’s own database, and — separately — one of the two receiving nodes left on the older build for the entire series. Pooled traffic routes through both nodes, so every number in the old table passed through a node that degraded as it received. We previously quoted the gain as +48%. We are withdrawing that percentage rather than restating it: the same curation rule we apply everywhere else disqualifies every pre-fix run, because each one routed through a node on the older build. That makes the 347.4 baseline void, and a percentage measured against a void baseline is not a measurement — it just looks like one. The corrected figures below stand on their own. We found all of this by re-checking our own rig, not because anyone challenged the numbers, and we are saying so here rather than quietly swapping the table.
| Run | Sustained pooled ingest |
|---|---|
| Sustained, run 1 | 529.5 instances/sec |
| Sustained, run 2 | 535.2 instances/sec |
| Sustained, run 3 | 536.5 instances/sec |
| Sustained, run 4 | 550.2 instances/sec |
| Fifth run (EXCLUDED — sender saw instances unacknowledged, see note) | 418.7 instances/sec |
We have removed the “Delivered 20,000 / 20,000” column that used to sit beside these rows. It read as a delivery ledger and it was not one: it was populated from sender-side counts. Our own reconciliation script documents the problem — the sender-side view cannot tell a lost instance from an unacknowledged one, and it marks a whole 50-instance batch failed on a single association fault. Receiver-side reconciliation exists, but it is a separate manual script and it was not run for these figures. So we do not have a delivery ledger for these runs, and we are not going to print a column that looks like one. What we can say is narrower and true: the router answered every one of those C-STORE operations with a success status.
Four clean runs, mean 537.9 instances/sec of pooled ingest. A fifth run is listed and excluded, because it is ours: at 418.7 it sat well below the other four, and we originally kept it. We had checked the wrong thing — every host showed 0.00% hypervisor steal and the balancer was at 99% CPU, so we concluded there was no fault to blame. The fault was not in the host telemetry. On that run the sender saw 19,657 of 20,007 instances acknowledged; 350 were not. Our curation rule says a run appears here only if the sender saw every instance acknowledged, and it did not. Excluding it moves this figure up, from 514.0 to 537.9 — we had been under-reporting ourselves by keeping a run we should have dropped.
These numbers replace lower ones we published earlier. The previous figures averaged 347.4 instances/sec. They were measured with a defect in our own software — an unbounded write-ahead log — and with one of the two receiving nodes left on the older build for the entire series, so pooled traffic ran through a node that degraded as it received. The improvement percentage we quoted alongside them is withdrawn: our own curation rule voids every pre-fix run, so there is no baseline left to measure an improvement against. The figures above are absolute, not relative to anything. We found this by re-checking our own rig, not because anyone challenged the numbers.
What pooling costs on delivery — and it costs less than we said
Azure figures, measured 24 August 2026 — after this correction pass had already begun, which is why this page said in three places, until now, that no such figure existed. Both routers 8 vCPU EPYC 9V74: identical silicon, so the two direct rates are comparable to each other and to the pooled one. All four Router-class hosts were reset to a clean database first, the sinks included. Published workload verbatim, three runs per condition, every run settling at exactly 20,000 delivered with zero failures.
This is the one correction on this page that runs in the product’s favour, so it gets the same prominence as the ones that do not. Everything we knew about the cost of pooling had been measured at the front door: a two-node pool accepts 0.45 of what the same two routers accept driven directly, and we had presented that as the price of availability. Measured on delivery, on runs that reproduced that same 0.45 at the front door, a pool clears 0.65 of the direct sum — and 1.24× what a single node clears on its own (delivery clock not recorded on that run — it predates the pinned measurement method). Pooling costs about a third of delivery throughput, not more than half. On the ingest framing a pool looked like a tax; on delivery it adds throughput as well.
| Measured | Router A, direct | Router B, direct | Sum of the two | Through the pool | Pooled ÷ sum |
|---|---|---|---|---|---|
| Ingest — accept + spool | 858 | 864 | 1,722 | 770 | 0.45 |
| Delivery — forwarded and acknowledged (delivery clock not recorded — predates the pinned §7.2 method) | 350 | 317 | 667 | 432 | 0.65 |
All rates are study-instances/sec.
The ingest row is why you should believe the delivery row. It is not there for symmetry: it reproduces the 0.45 pooling ratio we already published at 8 vCPU, and it does so on the same runs that produced the delivery figure. A favourable new number that arrives on a rig which no longer reproduces the unfavourable old one is not a measurement, it is a change of subject. This one reproduces it. And the pool split the work evenly to within 1% — 10,100 / 9,900 and 10,000 / 10,000 across the two nodes — so nothing here is an artefact of a lopsided balancer.
Why the two rows disagree. Ingest is the front door, and the extra hop through the balancer costs real acceptance capacity — every object is received once by the pool and again by a router. Delivery is the pipeline, and the pipeline is where a second node actually helps: each router forwards through its own outbound leg to its own sink, so the two legs run in parallel where the single front door cannot. Measuring only the front door was hiding the benefit the pool exists to provide. The published 0.45 understates what a pool is worth; it is a true ingest ratio and a misleading summary of pooling.
What this is not: a delivery figure for any other pool size, cloud or workload. It is one tier — 8 vCPU — on one rig. The pooled ingest ratios of 0.36 / 0.43 / 0.45 further down remain ingest ratios and are not restated by this result, and we are not going to interpolate a delivery curve through a single point.
Does a bigger load balancer help?
Azure figures. Only that rig answers this question: every Azure node runs the same processor model, so 2 → 4 → 8 vCPU changes core count and nothing else. Our Vultr rig mixes EPYC-Rome and EPYC-Milan across tiers, so its curve cannot separate core count from processor generation.
Up to a point, and then it stops. Quadrupling the load balancer from 2 to 8 vCPU buys 25% of ingest, and most of that arrives by 4 vCPU. Past there the constraint moves off the balancer and onto the routers behind it — measured directly: at 2 vCPU the balancer runs at 97% CPU while the routers sit at 73–76%; at 8 vCPU the balancer drops to 28% and the routers become the busiest thing in the path.
| Load balancer | Pooled ingest | vs 2 vCPU |
|---|---|---|
| 2 vCPU | 626.1 instances/sec | baseline |
| 4 vCPU | 741.2 instances/sec | +18% |
| 8 vCPU | 785.1 instances/sec | +25% |
Only one of our two rigs is quoted here. The other mixes AMD EPYC-Rome and EPYC-Milan processors across its tiers, so its curve varies core count and processor generation at the same time and cannot answer this question — however reasonable its numbers look. The rig quoted runs an identical processor on every node.
The August 2026 ingest-versus-delivery correction does not disturb this comparison. Every figure in it is an ingest rate, on both sides of every ratio, measured the same way — so the shape of the curve, and the conclusion that the constraint moves off the balancer by 4 vCPU, stand as measured. Only the column heading needed the word ingest. What the curve does not tell you is how the pool's delivery rate scales with balancer size. We now have pooled delivery at one tier only — 8 vCPU, in the section above — and one point is not a curve.
Where the router stops accepting images
Every store-and-forward router has a volume at which its operational database becomes the limit. We measured ours rather than leaving you to find it: eight sustained soaks, four per cloud, each pushed until the router began refusing images.
| Rig | Database size at first refusal | Individual soaks |
|---|---|---|
| Azure, 4 soaks | 486 MB mean | 453 / 458 / 462 / 570 MB |
| Vultr, 4 soaks | 499 MB mean | 438 / 463 / 514 / 583 MB |
The two clouds differ by 3% while runs on the same cloud vary by up to 33% — so the limit is a property of the software, not of the hardware you run it on, and it travels. That is the opposite of the rate figures above, which are firmly a property of your hardware. Both are worth knowing, for different reasons.
This one is unaffected by the ingest-versus-delivery correction above, because it is not a rate at all. It is a database size at the moment the router first refused an image — a threshold, measured in megabytes, on a clock-independent event.
Two things follow that we would rather say plainly. There is no gradual warning: ingest held between 66% and 103% of baseline right up to the burst in which refusals began — one run was measurably faster than its own baseline as it hit the wall. And the fix is retention, not a bigger disk. A retention policy that deletes nothing under load fails the same way on any database engine, just later.
The routers behind the pool, measured individually — ingest and delivery
This section previously described these figures as real end-to-end work: receive, write to spool, evaluate routing rules, forward by C-STORE to a live destination that drains. Three of those four stages were in the measurement. The fourth was not. The harness stops its clock at the sender's C-STORE response, and the router answers success as soon as it has taken the object onto its processing queue — forwarding runs asynchronously behind that queue, after the stopwatch has stopped. The 527 and 817 we published were ingest.
We re-measured both boxes on 24 August 2026, running the published workload verbatim — 15-second ramp, 500-instance warmup, 32 concurrent senders — three times per host, on a clean database, with the queues verified empty at settle and no failures in any run. Ingest reproduced the published figures. Delivery, counted as the objects the receiving node actually acknowledged and clocked from first arrival to last delivery, is the second number in the table below. Both hosts are the same cloud plan; the silicon we were actually given differs, which is exactly why we quote the clock with every figure rather than the plan name.
What the destination was, because it changes how you should read these numbers. Each router forwarded to a separate receiving node that accepted and wrote every object — four nodes per cloud, in two fixed pairs, rather than two hosts trading roles. The receiving nodes are not clones of the routers: on Vultr they carry 15 GB against the routers' 62 GB, and on both clouds they run at a lower clock. They were sized to drain without becoming the constraint, not to mirror the sender. That choice matters, and the alternatives we rejected show why. With a discard rule and no writes at all the same host reported 1,653/sec. With no reachable destination it reported 501/sec — back-pressured, its outbound queue simply filling.
We used to say the draining configuration was the one that “measures the product doing its actual job.” That claim is withdrawn. All three of those figures are acceptance rates — the same clock, stopped at the same place. What the spread between them actually shows is that acceptance is coupled to the state of the outbound queue: back-pressure downstream slows the front door. That is a real and useful property, and it is why we still publish the draining configuration rather than the discard one. But a number that moves with downstream conditions is not thereby a measurement of the forward. The forward is the delivery column, and it took a separate measurement to get it.
The consequence, stated plainly: a router sink is a cheap drain. It writes no archive catalog row to a database per object. Point the same router at a real archive on a relational database and expect the archive, not the router, to become the first thing that saturates. These are router figures, not whole-system figures — and that applies to the delivery column as much as the ingest one. Delivery against a cheap sink is the ceiling, not the expectation.
| Router host | Ingest — accept + spool | Delivery — forwarded and acknowledged | CPU at peak |
|---|---|---|---|
| 8 vCPU EPYC-Rome @ 2.0 GHz / 64 GB | 511 instances/sec 488–547 · published as 527 | 159 instances/sec 154–164 | 89.8% peak |
| 8 vCPU EPYC-Milan @ 3.25 GHz / 64 GB | 751 instances/sec 724–767 · published as 817 | 196 instances/sec 195–197 | 92.1% peak |
Ranges are the spread across three runs per host. The CPU column is from the original ingest campaign. Read the two rate columns as a pair: the Rome box accepts at 511/sec and clears at 159/sec, the Milan box accepts at 751/sec and clears at 196/sec. At the instant the published clock stopped, 8.0% of the workload had been delivered on the Rome box and 2.0% on the Milan box; the rest was still in the queue.
- Published ÷ delivery — 3.31× and 4.17×
- 527 against 159, and 817 against 196. This is how far the numbers we published overstated the forwarding rate, and it is the figure to use when correcting a document written off the old page — because the number in that document is the published one.
- Measured ingest ÷ delivery — 3.21× and 3.83×
- 511 against 159, and 751 against 196. This is Router's own front-door-to-pipeline ratio: a property of the product, measured on one clean campaign, and the right figure when the question is how acceptance relates to delivery rather than how wrong we were.
They differ because re-measured ingest came in slightly below what we had published (511 against 527, 751 against 817). We had been printing the smaller pair against the larger claim — the correction itself understated the error. Neither ratio is a substitute for the other and neither appears here without its numerator named.
Milan out-ingests Rome 751 to 511/sec but delivers only 1.23× faster (196 to 159). The forwarding leg does not scale with the processor the way the front door does. If your constraint is a backlog that has to drain, a faster processor buys you materially less than the ingest numbers imply — which is precisely the mistake this page was inviting when it published only one of the two.
One control run, reported because it is interesting and flagged because it is not a configuration you would deploy. On an Azure host — 8 vCPU EPYC 9V74 — with the spool placed on a RAM-backed filesystem rather than disk, delivery measured 336 instances/sec (four runs, 329–351) at association caps of 40/40/40, and 321/sec (three runs, 316–329) at caps of 32/10/3. The forwarding leg alone ran at 524 and 480/sec respectively, against ingest of roughly 724–755/sec; 19.4% of the workload had been delivered when the send-loop clock stopped. We ran it to see how much of the delivery gap is spool I/O. It is a deliberate control with the durability taken out, not a customer configuration, and we are not quoting it as a product figure.
Method: DOC-2026-490, the same controlled test method as every other figure on this page, with the delivery measurement added 24 August 2026. Raw harness output, the per-instance delivery counts and per-second CPU samples for each run are retained and available with the method on request. This campaign is one location of a planned multi-site series — the same rig, workload and method repeated on other clouds and on physical hardware, published here as each completes.
Five findings, including our own bottleneck
A benchmark that only produces flattering results is a brochure. These are the conclusions that changed how we size deployments — several of them inconvenient.
- For routine imaging ingest, cores mattered more than IOPS.
- On our test hosts — NVMe-backed instances with storage headroom to spare — quadrupling provisioned disk IOPS produced no measurable ingest gain for routine 512x512 CT. That workload was CPU-bound, not disk-bound. We are narrowing this twice over. It was measured on fast local NVMe, so it does not license the conclusion that disk speed is irrelevant on network-attached or lower-tier cloud storage, which we have not tested. And it was measured against ingest only — we have not run the IOPS ladder against delivery, so this finding says nothing about how disk affects the rate at which a backlog drains. Size CPU first; do not starve the disk.
- On ordinary cloud managed disks, storage binds long before CPU.
- The August 2026 Azure campaign put numbers on the narrowing above. The router wrote a measured 2,228 KB to disk per 512 KB instance delivered — each object was serialised at four pipeline stages — so sustained throughput needs roughly 2.2 MB/s of disk bandwidth per instance/sec. Measured against the storage tiers on that rig: the network-attached managed disk (163 MB/s, consistent with Azure’s P20 tier though the tier itself was not verified) projected a ceiling near 75 instances/sec, local temporary NVMe (817 MB/s) near 375, and only a RAM-backed spool cleared the Vultr-class rates above — which is a benchmark configuration, not a deployment. We then measured the managed-disk ceiling end to end: 77.8 instances/sec delivered, on both of two replicates, with the spool on that disk — within 4% of the projection, and slightly above it. The same host, build and workload with a RAM-backed spool delivered 337 instances/sec: a 4.3x storage penalty on delivery, measured, not modeled. Then we fixed the write path. Router 1.2.3.366 replaces two of the four per-stage disk writes with atomic renames; measured under the controlled method, write amplification fell from 4.26x to 2.33x (three replicates each, rate-matched within 0.4%) — 45% fewer bytes written per delivered instance — and the same managed-disk delivery ceiling rose from 77.8 to 140.4 instances/sec (three replicates, spread 0.3), a 1.80x delivery gain on identical hardware. One structural change rides along: delivery now runs at 0.91x the acceptance rate on that disk, up from 0.60x — so the backlog a node buffers mid-burst, and with it the exposure on a node loss, shrinks. These are bandwidth results; storage capacity planning is a separate axis and is unchanged by this work. If your spool lives on network-attached or lower-tier cloud storage, size the disk bandwidth, not just the cores — and see the delivery finding above for what changed in Router 1.2.3.366, measured rather than projected.
- The bottleneck moves with the workload shape.
- Every figure above was measured the way benchmarks usually are: one sender, one modality, one image size. A hospital is not that, so we built the realistic shape — five concurrent senders on separate hosts, five modalities from 64x64 SR objects to 2048x2048 DX (a 32x pixel-area spread), mixed association behaviour including one-instance-per-association senders. Under that mix the same router on the same disk is CPU-bound at 99%, delivering 102.9 instances/sec single-destination — where uniform CT on identical hardware is storage-bound at 140.4. A benchmark showing only uniform CT is not showing the constraint a real site hits, which is why we publish both. Integrity held under the hardest shape we have run: 10,600 offered, 10,600 accepted, zero refused, delivery receipts exactly one or two per instance on every replicate — no loss, no duplication, with interleaved multi-source studies at full CPU.
- A second destination costs 41% of instance throughput and ~3x peak memory.
- Fanning the realistic mix out to two destinations took instance throughput from 102.9 to 60.2/sec while total delivery work rose 1.17x — and peak process memory rose from ~1.0 GB to 2.4–3.2 GB, the largest memory effect we have measured. Two things to carry into sizing: count DELIVERIES when you plan a fan-out (10,600 instances to two destinations is 21,200 deliveries — reading receipts as instances overstates throughput by exactly 2x), and budget memory per destination, not per router. Pooled on the current build: two nodes clear 244–250 instances/sec, 1.75x a single node — down from 1.96x on the prior build only because the single-node ceiling rose from 77.8 to 140.4; the pool did not get slower, the baseline got faster. And for sizing a pool member to survive its partner: a node at pooled (roughly half) load ran a measured 746 MB mean; the same node absorbing the whole load ran 1,159 MB — a 1.55x surge ratio, sub-linear because fixed overhead does not double. That surge multiplier and the per-destination fan-out multiplier above have different causes and do not compose — size for each separately.
- The inbound-connection limit binds before the CPU saturates.
- Our own bottleneck, and it is a configuration ceiling rather than a hardware one — an ingest-side finding from small-object load, where rates run high enough for the association cap to govern. On real 512 KB CT the storage write path binds first at ~155 instances/sec (the RI-2026-644 ladder), well below where the cap engages. Worth knowing before you size a box around CPU headroom you will not reach.
- High memory use under load is by design, not distress.
- The router deliberately uses up to ~75% of host RAM as a managed ceiling and recovers fully afterward — confirmed by direct measurement 2026-08-25: the configured heap ceiling truncates the same 3D workload at ~11 GB under a 16 GB cap, 21–24 GB under 32 GB, and 41–43 GB uncapped on 62 GB, accepted in full at every tier; removing the cap nearly doubled peak use at identical load, proving the ceiling was what bounded it. A monitoring alert on "memory high" here is a false positive.
- The 3D mammography memory ceiling scales with the host.
- The 75% heap ceiling adapts to the box: smaller hosts run the same concurrent-volume workload in proportionally less memory, at the cost of latency rather than refusals in our re-measured runs. There is no single answer to "what limits mammo" — it depends on the box, which is why we measured the same workload across memory tiers.
- Capacity grows by adding nodes, not bigger boxes.
- Past the comfortable band, the scale-out path is a SynthIQ pool that load-balances across routers and places large studies on nodes with memory to spare — not a larger single host. The delivery measurements sharpen this rather than soften it: across the two routers we measured both ways, the faster box led on ingest by 751 to 511/sec but on delivery by only 1.23x (196 to 159). The forwarding leg does not follow processor speed the way the front door does. Measuring the pool itself on delivery makes the same point from the other side: a two-node pool clears 1.24x what a single node clears, where on ingest alone pooling had looked like a tax paid for availability. Two legs forwarding in parallel is a real gain that the front-door measurement could not see.
How we measured
The rig
One VM generated DICOM load; a separate VM ran the system under test — a single router instance, not a pool — and they communicated over a private network with nothing else competing for resources. The system under test was measured at three host sizes — 4 vCPU / 16 GB, 4 vCPU / 32 GB and 8 vCPU / 64 GB — so the limits could be watched moving with the box rather than read from a single data point. Generator utilisation was recorded alongside every result, so a figure is only reported where the generator was demonstrably idle and the constraint was genuinely in the system under test.
The corpus
Synthetic DICOM throughout — no patient data of any kind. The harness emits studies from five modality emitters (CT, MR, CR, US and MG) rather than replaying a clinical archive, which is what makes the corpus profile publishable and the runs repeatable. The stress cases are deliberately the two opposite shapes: few enormous multi-frame objects (tomosynthesis) and very many small objects (a 10,000-slice CT).
The harness
A self-contained .NET 8 console application. It provisions its own sources, destinations and routing rules; emits the synthetic studies through the product under test; sweeps a rate ladder recording p50 / p95 / p99 latency; and tears down everything it created on exit. Production wiring is never touched — only the artifacts the harness itself created are removed.
We have deleted a sentence that used to sit in this paragraph. It said the harness “verifies every study actually landed by querying the destination over DICOMweb QIDO-RS rather than trusting a send count.” That was not true of the harness that produced the figures on this page. We maintain two load harnesses; the QIDO-RS verifier is in the other one, which our own run plan says is not the capacity tool. The harness behind every published figure here contains no QIDO code at all — its only pre-run check against the destination is a DICOM C-ECHO, and it trusts exactly the send count the deleted sentence disclaimed. This was a claim about the rigour of a measurement that was not true of that measurement, so it is removed rather than reworded.
What the harness actually records, stated plainly: it counts sender-side C-STORE successes and divides by the elapsed time of the send loop. That is the number, and it is the only timing mode either harness has. Because the router returns success as soon as it has queued an object for processing, this clock measures ingest. It stops before forwarding completes.
The delivery figures were produced separately and by a different count: the objects present in the receiving node's sent/ directory — one file per instance the router forwarded and the sink acknowledged — clocked from first arrival to last delivery, with the queues verified empty at settle before the run was accepted. Three runs per host, on a database reset to empty beforehand.
If you are evaluating us and want to run it in your own environment against your own hardware, ask — we would rather you measured us than took our word for it.
A variable we did not control
Router ingest declines as its operational database grows — that is the volume limit measured further up this page, and it is the one condition on this rig capable of moving a rate without anyone touching the workload. Our soak script knows this and handles it: it resets the operational database before each measurement, and says why in its own comments — the sink is itself a router and accumulates as it receives, so leaving it loaded would mean the measurement degrading for a reason that is not the variable under test.
The campaign scripts that produced every published figure on this page do not do that. They never reset the database and never record its size. And every host on both rigs was found above the lowest database size at which we measured a refusal (438 MB) — the two busiest far above the top of that range, at 1,467 MB on the Rome router and 1,901 MB on one of the sinks. So the published figures were taken at an unknown and drifting database volume, on a variable we had already demonstrated matters. We are disclosing it rather than restating the figures, because we do not know the size at each run and will not reconstruct one. The 24 August delivery measurements were taken on a clean database; the ingest figures beside them were not.
DOC-2026-490 — Performance Benchmark Test Method: DICOM Load/Verification Harness. It states the corpus, the rig, the isolation discipline, how ingest and delivery are each counted, every measurement recorded, and the boundaries of what the method does not establish. Ask for it by number, along with the harness itself, and we will send both — to customers evaluating us, and to vendors who want to check our figures.
Request DOC-2026-490 and the harnessWhat we have not measured
Delivery, on most of this page. It has been measured on exactly two single hosts — the two Vultr routers in the table above — on a two-node Azure pool at one balancer tier, and on one deliberate control on Azure with the spool on a RAM-backed filesystem. It has not been measured for either host in the routine-imaging table, for the 3D mammography concurrency ladder, or at more than that single point on the load-balancer scaling curve. Every rate in those places is ingest. We are not going to derive a delivery figure for them from the handful we have; the ratio between ingest and delivery already differs by processor generation on the only two hosts where we can compare, which is reason enough not to extrapolate it onto a third.
End-to-end AI-vendor egress latency — the added round trip through the cloud routing tier to an AI vendor and back — has not yet been measured and published. The design target is <200 ms p99 added latency for radiology-report-sized payloads, but a target is not a measurement and we will not quote it as one. It appears here when the run is done.
The pool scaling factor — pooled ingest divided by the sum of the same routers' ingest measured individually — is now published: it is Figure 2 above, and it runs 36% / 43% / 45% across the three Azure tiers. Both sides of that ratio are ingest rates measured the same way, so the ratio holds; what it does not describe is the pooling cost on delivery. That is a separate measurement, now made at the 8 vCPU tier and reported under What pooling costs on delivery above, where the pool comes out at 0.65 on delivery against 0.45 on ingest — cheaper on the pipeline than at the front door, and not a figure the ingest ratio could have stood in for. Getting there found a bottleneck on our side rather than a limit of pooling: the pool tier was opening a new DICOM association for every single image it forwarded, instead of reusing one across many. Each router was paying the full cost of setting up a connection per image and spending its processor on handshakes rather than on images.
We have fixed it, and the figure gets published once the run is repeated against the corrected build. We are saying this here rather than quietly republishing a better number later, because a benchmark page that only ever reports flattering results is not measurement, it is advertising — and the first thing this campaign measured carefully turned out to be a defect of our own.
The figures
Two results carry most of the argument, and both read better drawn than tabulated: the volume limit is about SPREAD, and the pooling cost is about a gap. Every value plots the measured runs, not a smoothed curve. The rates plotted in the second figure are ingest rates, as is everything they are compared against.
Each dot is one sustained soak pushed until the router began refusing. The spread within a cloud is far wider than the gap between them.
Means 486 and 499 MB — 3% apart, on hosts whose raw ingest rate differs by 60%. The limit follows the software, not the hardware.
Same two routers, same workload. Measured individually, then driven through the load balancer. Both bars are ingest — instances accepted at the door and spooled, not instances delivered. Figure 4 measures the difference.
Pooling costs 55–64% of the arithmetic sum of the two routers' ingest rates. A bigger balancer recovers some of it — and stops helping once the routers behind it become the constraint. These percentages survive the 2026-08-24 correction: both sides of the comparison are the same kind of number, measured the same way, so the ratio is unaffected. Only what the bars are called has changed.
Identical build, identical workload, identical topology. The only difference is the hardware underneath. As in Figure 2, both series are ingest acceptance rates, not delivery rates.
The gap runs 1.16× to 1.35× — compare Figure 1, where the volume limit differs by 1.03× between the same two clouds. The limit follows the software; the ingest rate follows your hardware.
One uncontrolled variable in Figures 2 and 3, disclosed here because it is not in the run plan: the soak behind Figure 1 resets the operational database before each measurement, and says why. The campaign runs behind these two figures never touch that database and never record its size — and every host on both rigs was found well past the refusal range Figure 1 plots (1,467 MB on one router, 1,901 MB on one sink). So these rates were taken at an unknown, drifting database volume.
Two single routers, not the pools in Figures 2 and 3. Published workload verbatim, re-run 2026-08-24: three runs a box, zero failures, queues verified empty at settle, clean database. The left bar reproduces the published measurement. The right bar counts one instance per file the destination acknowledged, clocked from first arrival to last delivery.
Ingest overstates delivery by 3.21× and 3.83× — the gap between what the published figures measure and what they were taken to mean. At the moment the published stopwatch stopped, 8.0% and 2.0% of the work had actually been delivered; the rest was still in the queue. Nothing was lost — received and sent matched exactly on every run — but it had not left yet.
Size on the delivery number, not the ingest one. Between these two boxes ingest rises 511 to 751 while delivery rises only 159 to 196 — 1.23×. The forwarding leg does not scale with the processor the way the front door does.
A deliberate control on Azure with the spool on tmpfs — not a customer configuration — delivered 336/sec at 40/40/40 association caps (four runs, 329–351) and 321/sec at 32/10/3 (three runs, 316–329), against 524 and 480 for the forwarding leg alone. Even on tmpfs, only 19.4% of the work had been delivered when the clock stopped.
What we discarded, and why
Anyone can put a number in a whitepaper. The campaign behind this page produced 251 measured runs, of which we discarded 165. Each discarded run is published with the reason it was dropped, because a table of only the surviving runs proves nothing about how they were chosen. That count was 160 until we audited our own curation and found five runs we had kept that our own acknowledgement rule already disqualified — on the largest of them the sender saw more than half the workload go unacknowledged, and it still carried no void reason.
- Runs measured before we found the unbounded write-ahead log in our own router. Ingest on those decayed with cumulative volume, so any rate from them is a function of how much had already passed through that node.
- Runs where one of the two receiving nodes had been left on the older build. Pooled traffic routes through both, so those numbers passed through a degrading node — including every pooled figure we had previously published.
- Runs on the rig whose processors differ between tiers, where a scaling question was being asked. The numbers are real; they just cannot answer that question.
Alongside the discards we publish the conditions: the processor model, core count and memory of all fourteen nodes, and the size of each router's operational database at the time of measurement — 477 MB to 1.9 GB, against refusal thresholds we measured at 438–583 MB. Those ranges overlap, and the top of the first sits well above the second: the database sizes were captured at end-of-series rather than per run, so we cannot claim every run was measured below the threshold, and we do not. Ingest declines as that database grows, which is why the condition is published alongside the rate rather than omitted. The campaign scripts never reset that database between runs and never recorded its size — see A variable we did not control in the method section above.
The raw result files, the per-instance send manifests, the receiver-side delivery counts from the 24 August runs, the CPU samples from every host, and the scripts that produced all of it are retained under version control and in our regulated artifact store, each file individually hashed. If a number here matters to your decision, ask us for the run behind it.
Full papers
Each number above comes from one of these, both maintained as controlled documents in the Synthology quality-management system.
- Router Pools: What Pooling Costs, What It Buys, and Where It Stops
The paper behind this page. Every figure above, with the method, the conditions each was measured under, and all 165 discarded runs with the reason each was dropped.
- SynthIQ Capacity: The Heaviest Studies on the Smallest Box
Memory behaviour under tomosynthesis and high-object-count CT on 2 vCPU / 8 GB.
- Right-Sizing DICOM Routing for Mammography
Router ingest, IOPS sensitivity, and the moving mammography bottleneck across three host sizes.
All software described here is general-purpose, non-device software under FD&C Act §520(o)(1)(D). Nothing on this page is a medical device or provides a clinical diagnostic function. Figures are measured on the hardware stated and will vary with your configuration, network, and study mix.