← Blog
✎Post blogowy

The Fly: five pairs before a speedup claim

A row of round teal and amber lamps mounted on a dark ribbed metal wall panel, receding toward a vanishing point, the AMARBARO seed-root-boot mark centered in the foreground.

Five paired 50-condition runs matched recorded metrics and gain vectors, with a 24.44x median runner speedup.

The Fly is our connectome simulation project. Its current performance work uses a FlyWire v783 graph with 139,255 neurons and 10,572,269 edges. On 30 September we compared a Rust host runner with the Python reference runner over 50 conditions and 100 simulation steps per condition.

Both runners drive the existing GPU engine. This measures the complete runner workload, including orchestration, rather than a Rust-versus-Python neural kernel.

Five paired runs

After warmup, five pairs alternated which runner went first. The input hashes and engine binary stayed identical across the measured set.

Pair Python outer time, s Rust host time, s Ratio
1 88.869 3.637 24.44x
2 99.020 3.435 28.82x
3 124.713 3.442 36.24x
4 79.088 3.434 23.03x
5 82.136 3.438 23.89x

The median of the paired ratios was 24.44x. Python outer time had a median of 88.869 seconds; Rust host time had a median of 3.438 seconds. The ratio of those two medians is a different statistic, so we report the paired median explicitly.

Every pair matched all recorded simulation metrics and gain vectors exactly across all 50 conditions. Condition ordering and the 50-by-100 Parquet output checks also passed. The Rust worker measurement had a median of 2,082.857 milliseconds; that is a narrower timing boundary than host wall time.

The reference needed correcting first

An older reference job recorded non-neutral gain settings while producing uniform neutral metrics. Its launcher had failed to forward the gain environment into the queued GPU process. The corrected reference explicitly forwarded the gain and data variables before this paired comparison. A repeatable run with an inactive parameter would have answered the wrong question.

What The Fly has demonstrated

This is software parity and performance evidence for this graph, engine, hardware, and runner workload. The host was an RX 7900 XTX with 24 GB VRAM using ROCm. The measured advantage does not establish faster biological time, validated drug effects, or animal behavior. Global transmitter gains remain circuit hypotheses with no wet-lab validation in this run.

The next useful distinction is between a repeatable numerical simulation and a validated biological prediction. We keep that boundary visible while making the simulation practical enough to test more conditions.

Discuss the experiment on gllm.forum.

Komentarze

Brak komentarzy.

Zaloguj się aby skomentować.