Skip to main content

Measuring mump2p on mainnet: earlier blocks on three clients, none of it significant

· 16 min read
Stefan Kobrc
Founder RockLogic
StereumLabs AI
Artificial Intelligence

We ran Optimum's mump2p gateway, version 1.1.1 at commit 66e391d, on 36 mainnet full nodes in a single datacenter for 11.5 days at the end of August, one gateway on every node, across the full six-by-six matrix of consensus against execution clients. Blocks arrived earlier on three of the six consensus clients. On a fourth about half the move is real, on a fifth the metric cannot say either way, and on the sixth the treated nodes got slower rather than faster. None of the numbers is statistically significant, and the reason is built into the design rather than into the data.

The apparent effect, before we looked hard, was 40 to 110 ms. We then spent a second pass attacking our own control group, because the treated nodes are bare metal in Vienna and the controls are GCP in Frankfurt, and a difference-in-differences absorbs a constant offset between two datacenters but never a drift. What survived that pass is smaller, is carried by fewer clients, and rests on a different kind of evidence than the control group provides. This post is about which parts held.

Read this first
  • This is not a verdict on mump2p. It is one measurement, on one fleet, for 11.5 days, with no validators attached. Where we found an effect we say how large and how much we trust it; where we could not, we say that too.
  • These numbers describe v1.1.1. We ran optimum-gateway v1.1.1 (commit 66e391d). Optimum is iterating quickly, and a newer release with better statistics is on the way, so read this as the build we ran at the end of August, not the current one. It is an early build, and too soon to say whether these numbers are the ceiling of what the technology can do.
  • Nothing here is significant at 0.05, and nothing here could have been. With k gateway-free rebuild boundaries to build a null from, the smallest empirical p this design can produce is 1/(k+1), which is 0.059 to 0.077 for our history. Four clients sit exactly on that floor. That means "as extreme as this design can resolve", not "proven".
  • Optimum's own published figure measures something else. Their roughly six-times number is average block propagation on the Hoodi testnet against an external baseline. Ours is a within-fleet, mainnet difference on 36 nodes. The two are not comparable, and we are not testing their claim.
  • Arrival is not usable arrival. On lodestar the gain is gone by the time the block becomes head, and on prysm the client's own processing time grew over the window. A block that lands earlier is not automatically attested earlier.
  • The strongest evidence does not use the control group at all. It comes from individual machines that crossed the fleet rebuild with unchanged IP, consensus build and execution build, and moved anyway.

What mump2p is

mump2p is Optimum's block-propagation layer. Optimum is a networking company built on random linear network coding, a coding scheme from MIT in which a message is split into fragments and sent as random linear combinations. Any sufficient set of coded shards reconstructs the message, and relays can re-encode without waiting for the whole thing. The optimum-gateway runs beside a consensus client as an ordinary libp2p peer: it subscribes to the beacon block and attestation gossip topics, forwards them into Optimum's coded mesh, and re-injects what that mesh delivers back into the node's local network. The premise is that the coded mesh reaches the node sooner than plain gossipsub does. It needs no client patch and runs alongside gossipsub rather than instead of it. We ran the gateway on every treated node and changed nothing else about the clients.

What we did

This is our standing measurement fleet, the same rig behind the slot-timing series. Every consensus client records, for itself, how late into the slot each block arrived. We took that metric as a difference-in-differences, which compares the change on the treated fleet against the change on an untreated control over the same period, so anything constant to both, like one datacenter simply being faster, cancels out. The treated fleet is in Vienna and the six control nodes are on GCP in Frankfurt; the gateway window (2026-08-27 to 09-07) is compared against the rebuild window just before it. Baseline arrival on this fleet is about 1,950 ms into the slot, so a real effect of a few tens of milliseconds is what we are trying to resolve.

The honest summary is one line per client.

ClientWhat we foundSurviving estimateTreatment-side
lighthouseearlier block, survives nine ways≈ −120 ms~70%
prysmearlier block, survives≈ −60 ms (gossip −78)~100%
grandineearlier, about halved from the first cut−31 to −52 ms~90%
lodestarabout half of it is control drift≈ −45 ms, half real~51%
tekunot a clean latency shift≈ −4 ms, indistinct from 0n/a
nimbusrefuted: no gain, all control drift+14 ms0%

Treatment-side is the share of each move that came from the treated nodes themselves rather than from the control drifting, so a higher number is stronger evidence.

Six consensus clients measuring the same mump2p intervention: lighthouse, prysm and grandine show blocks arriving earlier, lodestar about half control drift, teku indistinguishable from zero and nimbus slower, four of the six sitting on the design's significance floor with teku and nimbus above it

The negative digit means earlier. Read the column of estimates and the story is consistent for the top three and falls apart below them, which is roughly what a first look at an effect this small would look like, if it is real, which a design this confounded cannot establish.

Why we do not fully believe the number

The control group does not hold cleanly on any client. We took the Vienna-minus-Frankfurt gap on every gateway-free rebuild window back to February, and it drifts: on lodestar by 2.47 ms per week (t = 3.31), on grandine and teku comparably. Worse for a clean story, every client has at least one gateway-free window with a gap as large as the treated one. lighthouse is the starkest, with gaps of 260 and 812 ms in early commissioning windows when the control VM was still warming up. So the claim "a gap this size needs a gateway" is simply not supportable from this history.

The lighthouse control gap across every rebuild window back to February, showing the treated window at −120 ms is the most extreme of the mature era but that gateway-free commissioning windows reached −260 and −812 ms

Two things keep the result standing despite that. The drift runs against the effect on every client that carries one, so detrending makes the estimate larger, not smaller; the gap is not being manufactured by the control sliding the convenient way. And the load-bearing evidence does not touch the control group.

The evidence that is not a datacenter artifact

On lighthouse, prysm, lodestar and grandine there are individual machines that crossed the fleet rebuild with the same IP, the same consensus build and the same execution build, and whose block arrival moved anyway. Differenced against their controls, that within-host move is −121 ms on lighthouse, −54 on prysm, −41 on lodestar, and +1.89 percentage points on grandine's bucket measure. Each figure rests on the two or three such hosts per client, so these are small-n contrasts, not fleet averages. On lighthouse the treated machines themselves moved about −87 ms, the largest step any of them took across any rebuild in the record. The controls behind these within-host figures moved +37 ms on lighthouse and only a few ms on prysm and lodestar. At the coarser cohort level lodestar's control drifts more, on the order of +15 to +22 ms across the boundary, which is why the table marks it only about half treatment-side.

Treated hosts that kept identical IP, consensus build and execution build across the boundary moved by −87, −54 and −38 ms while their controls moved +37, +3 and +3 ms, giving within-host differences of −121, −54 and −41 ms

A move on a machine whose IP, consensus build and execution build did not change cannot be an artifact of one datacenter being faster than another. It does not, on its own, separate the gateway from the fleet rebuild that happened at the same instant; the placebo test over gateway-free rebuilds is what speaks to that. Together they are why we treat the three-client result as a signal worth publishing, not a proof, from a design this compromised.

The nimbus result is the counterweight. The apparent gain there was entirely a single GCP control degrading by 58 ms over the same window; not one nimbus host sped up, and the treated side, if anything, drifted 14 ms slower. When your control is one VM, "the control had a slow week" and "the treated fleet had a fast week" are not separable, and on nimbus the evidence points squarely at the former. Read it as no detectable benefit, not as mump2p harming the client.

Why none of it is significant

The design gives itself away here. With k gateway-free boundaries to draw a placebo distribution from, the smallest empirical p it can report is 1/(k+1). For our record that floor is between 0.059 and 0.077, and four of the six clients land exactly on it. A true effect of any size could not have produced a smaller p. On top of that there is a single control host per client, and at lighthouse that one VM carries more than half of the placebo noise. And there was exactly one treated boundary with no period after it, because the fleet was torn down on 2026-09-08, so the strongest test we can imagine, whether the gap snaps back when the gateway is removed, is permanently out of reach.

Where this sits next to Optimum's own numbers

Optimum's headline figure is roughly six times faster: about 150 ms average block propagation on the Hoodi testnet against a ~1 s gossipsub baseline drawn from ethPandaOps monitoring, across their own 30-node gateway fleet. That is an absolute arrival time on a testnet measured against an external baseline, which is a different quantity from a marginal, within-fleet difference on mainnet, and we are not testing it. An economics writeup on ethresear.ch co-authored by Optimum's Muriel Medard assumes a 50 to 150 ms improvement and maps it to roughly a 0.66 to 1.97 percent increase in validator revenue for large operators. That is a relative increase in revenue, not points added to APR; on a 3 percent base it is a few hundredths of a point. Its advantage metric is the P80 propagation latency relative to the four-second deadline, a tail quantity that our median-level arrival numbers do not isolate. Our surviving numbers span the low-to-mid part of that millisecond window: grandine and lodestar near the 50 ms floor, prysm in the middle, and lighthouse's roughly 120 ms near its top.

There is a structural reason to expect a bounded gain, and it is our reading of gossipsub rather than an Optimum claim or a spec guarantee. A node acts on the first copy of a block it receives, so an extra source only helps when it beats every peer already delivering that block. That caps the gain at how much the gateway beats the fastest peer already in the mesh, not at the whole propagation time, whether or not the mesh is full. On our fleet the block-mesh peer count barely moved when the gateway arrived, from 6.91 to 6.96 against a target of eight, which suggests it rarely became a mesh peer at all; an operator can likely do better by attaching it as an explicit direct peer, which we did not do and which is the obvious next experiment. None of this is a limit Optimum imposes: they present mump2p as opt-in and additive, run alongside gossipsub, with an individual-node benefit that does not require network-wide adoption. Our narrower point is that 36 gateways among thousands of mainnet nodes cannot move the network-wide propagation number itself.

This is, as far as we can find, an early independent measurement of mump2p on mainnet. The closest prior work is the Lido PERCH pilot thread, where node operators reported testnet results (HashKey Cloud saw roughly 200 ms lower attestation propagation), and FP-Validated's open-source attestation-forensics tool, a measurement tool rather than a published result. Both are worth reading alongside this.

What we could not test

Arrival is not the same as a usable block, in several ways this measurement cannot see past. On lodestar the head-import metric shows the gain gone by the time the block becomes head, the point the client's fork choice adopts it as the chain tip, and on prysm the client's own processing time grew by 11 to 46 ms over the window while staying flat on the controls, so an earlier arrival there is partly eaten before the block is usable. Since Dencun a block is not attestable until its blob sidecars arrive and pass a data-availability check, and the gateway as we ran it carries beacon block and attestation gossip, not the blob topics, so an earlier block can still wait on blobs. Much of the ~1,950 ms baseline is proposers publishing late to capture MEV rather than propagation time, so the slice mump2p can move is a small, high-variance part of the total. We report central tendency, not the tail, and it is the tail that governs missed attestations: our slot-timing series found that on this fleet almost none of the four-second deadline pressure comes from blocks arriving late, 0.11 to 0.16 percent, and nearly all of it from processing after they arrive, 5.1 to 9.0 percent, which an arrival gain does not touch. And these were full nodes with no validators attached: attestation gossip still crossed them, but the validator's own attestation-timing and reward path was never exercised. We also measured only blocks arriving at our nodes, not the propagation of the blocks a node publishes outward into the network, where a gateway may help differently and which, with no validators, we had no block to test. We did not profile the gateway's own bandwidth and CPU footprint here either, though the fleet exposes those through the same dashboards this series is built on. Anyone evaluating mump2p for attestation or reward performance will need validators attached first, and should treat the numbers here as a measurement of block arrival and nothing more.

Methodology

The treated fleet is the NDC2 deployment in Vienna, 36 bare-metal hosts on identical 12-core machines, one Optimum gateway per host, running the full six-by-six matrix of consensus and execution clients. Controls are six GCP europe-west3 VMs that never carried a gateway. The window is 2026-08-27 16:05 to 09-07 23:17 UTC; the gateway fleet was decommissioned 2026-09-08.

  • The metric is each consensus client's own block-arrival timer: beacon_block_delay_gossip (lighthouse), block_arrival_latency_milliseconds and its gossip counterpart (prysm), beacon_block_gossip_slot_start_delay_time (grandine), lodestar_gossip_block_elapsed_time_till_received (lodestar), beacon_block_import_delay_counter_total{stage="arrival"} (teku), and beacon_block_delay (nimbus). teku and grandine expose no _sum, so their primary measure is a bucket-crossing fraction rather than a mean, and every millisecond figure quoted for them is model-derived and flagged as such. Each client measures arrival on its own instrument, which is why we compare a client only against itself and never pool the six; our instrument reference documents those differences in full.
  • The estimator is a difference-in-differences of the Vienna-median-minus-Frankfurt gap, treated window against the prior rebuild window, cross-checked against a within-host contrast on machines whose IP, consensus build and execution build were unchanged across the boundary.
  • Significance is by placebo randomisation over every gateway-free rebuild boundary back to 2026-02-18. With k such boundaries the empirical-p floor is 1/(k+1), which is 0.059 (lighthouse, grandine, nimbus), 0.062 (prysm), 0.067 (teku) and 0.077 (lodestar). Each analysis was re-checked adversarially; where the recheck overturned a figure, the corrected one is used here.
  • Known confounds, unresolved. The gateway arrived at the same moment as a fleet rebuild and an execution-client bump; each control is a single VM; the lighthouse series required a scrape-phase bias correction before its effect was separable from noise; and there is no post-teardown period, so no snap-back test exists.
  • Fairness note. The build we ran, v1.1.1 (commit 66e391d), shipped under Optimum's reference-only Microsoft Reference Source License; the project was relicensed to MIT the same day, 2026-08-07, so current builds are open source but the tree we measured was not itself an open-source release. Optimum's performance and economics figures cited above are the company's own, on testnet or in controlled clusters, and are linked in place.

Coming next

This is a single measurement of one build, and it wants a rerun on the newer gateway version, with validators attached, a longer window, and an untreated arm inside the same datacenter, which is what would turn a within-host contrast into a real control. If you run Ethereum infrastructure and want this lens on your own nodes, the queries and fleet labels behind this post are documented in build your own dashboards, the stack itself in the measurement stack we described here, and you can reach us at stereumlabs.com or contact@stereumlabs.com.