Skip to main content
Back to posts

🔧 The slowest part of my speed comparison

niklas-heer/speed-comparison

A repo which compares the speed of different programming languages.

749 113 Python MIT Updated 1 days ago

My speed-comparison repository runs the same calculation in as many programming languages as I can get installed. What people come for is the chart.

What I have actually spent most of my time on is installing programming languages.

Compilers, interpreters, libraries, base images, version checks, build flags. Before anything can race, somebody has to get all the contestants to the starting line, and over the years that turned into the more interesting problem. The project has been through three approaches to it. Reading the commits is a good reminder that a tidy chart can have a very untidy cupboard behind it.

THREE WAYS TO GET TO THE STARTING LINE

The benchmark stayed. The plumbing changed.

pacman -S pypy ruby php rust go nodejs …

Every language moves into the same container. The environment is a shared maintenance concern before the benchmark starts.

A map of the repository's evolution, not a timing comparison. The linked commits below show the actual configurations.

2018: one Arch container with everything in it

In February 2018 I moved the environment to Arch Linux. The Dockerfile installed the base tools, then a long list of language runtimes with pacman, then the plotting tools, the Python dependencies and the package helper.

The idea was simple. Build one big image with everything in it, run the comparison inside.

That worked until the image itself became the thing I was maintaining. Every new language was another package in the shared home. Updating the distribution and updating the contestants were nearly the same operation, and not one I looked forward to. I wanted a benchmark suite and was slowly acquiring the responsibilities of a small Linux distribution.

2022: an Earthly target per language

The first Earthfile arrived in October 2022. Splitting the work into targets helped. The shared preparation and the benchmark steps could be reused, and each language got its own instructions where it needed them.

There was a short NixOS detour in there too. The very next commit removed those images again because of the load they put on the machine. The road to declarative infrastructure had a round trip in it.

The bigger problem stayed. The recipe for “install this language” still depended on where the language came from. One target used an Alpine package. Another needed a Debian-based image. A third started from the language’s official container. Bumping a version could mean changing a package name, swapping an image tag, editing an install script, or finding out that the new image brought a different version of something else along.

The Python/Numba addition is a typical case: a python:3.13-slim base, Debian preparation, compiler and development packages, then pip install numba. Perfectly reasonable. Also one more special environment to understand the next time anything changes.

None of this is Earthly’s fault. Earthly can reuse targets and run work in parallel. What had gone wrong was that this project had collected a different installation recipe per language, and the steps for preparing toolchains, building programs and collecting results had grown into each other.

2025: describe each language once

In December 2025 I added a Dagger pipeline. Dagger does the container work, driven from Python. Nix packages, managed through Devbox, supply the language environments.

The part I like most is embarrassingly small. It is a Python dictionary.

dagger-poc/languages.py holds the LANGUAGES definitions. This is the real Go entry with the display metadata left out:

"go": Language(
name="Go",
nixpkgs=("go@1.25.4",),
file="leibniz.go",
compile="go build -ldflags='-s -w' -o leibniz leibniz.go",
run="./leibniz",
version_cmd="go version",
),

Package version, source file, compile command, run command and version check all sit together. The same data drives the build, the benchmark, the validation and the update tooling.

For an ordinary version bump I change the declaration and let the pipeline do the rest. I no longer have to remember which Docker image happened to contain the version I wanted. The version-check workflow is built on the same definitions. There is something very satisfying about declarative configuration when the alternative is rediscovering your own installation instructions every few months.

Build the toolchains before the race

The new pipeline builds the language environments ahead of time and stores them in the container registry. The benchmark source is added later, when it is needed. That gives me a clean line between preparing a toolchain and running a measurement.

The commits show that line getting sharper:

Preparation can now be split into independent jobs. The timed runs are a different matter. I do not want every benchmark competing for the same CPU at once, so the runner still works through its selected targets one at a time.

What still needed fixing by hand

The migration has plenty of repair commits too. Swift wanted particular Nix packages and headers. Commands had to go through devbox run to see the right environment. A few languages still need explicit setup steps beyond a package declaration.

The Earthfile is also still in the repository, next to a directory that is still called dagger-poc. This is a migration in progress, not a claim that the old system disappeared overnight.

What did improve is where the knowledge lives. Versions and exceptions have one representation I can test. Building environments and running benchmarks have a clear boundary between them. Updating the collection is no longer an expedition through unrelated containers.

The chart is still what most visitors come for. The work I am happiest about is the part that makes the next update boring.