The icechunk to zarr conversion finished, although not when I said it had; benchmarking Tessera inference on Intel’s AMX silicon; and day10 spreading to four architectures while finding real bugs.
Tessera #
The conversion finished #
The whole conversion took about 13.5 days of wall clock from the first job on the 13th, roughly 125,000 vCPU-hours, and cost about $3,043 on Fargate Spot. The same compute on demand would have been around $9,200, so Spot saved something like $6,100. Storage is billed to source.coop under the AWS Open Data programme rather than to us.
The final counts, per year:
| Year | Shards |
|---|---|
| 2017 | 93,915 |
| 2018 | 97,652 |
| 2019 | 97,723 |
| 2020 | 97,887 |
| 2021 | 97,902 |
| 2022 | 98,091 |
| 2023 | 98,094 |
| 2024 | 98,093 |
| 2025 | 98,068 |
2017 has around 4,000 fewer shards than the other years, and in the collage above it is the only panel with red in it. The red indicates that there is data for that shard in other years but not this one. dClimate applied a quality threshold to avoid processing tiles with sparse data and there are few observations of these areas in 2017.
Coverage Tool #
Annoyingly, I declared 2025 finished on Tuesday. Anil ran the RGB pyramid and found gaps in the data. The problem was that I had been trusting Fargate Batch exit status and for all of 2025, they reported SUCCEEDED.
Before I could trust the exit status again I needed to see what data was in the Zarr store. Previously, I’d use mtelvers/tessera-mosaic for continent scale images which did partial reads against the NPY files on S3. Since I already had mtelvers/ocaml-s3 and mtelvers/ocaml-zarr I could relatively easily generate a map directly from the zarr store.
The tool initially just coloured each pixel by whether a shard object exists and does the whole world in about 75 seconds with roughly a thousand list requests, anonymously, using no credentials at all.
However, I wanted to actually see “data”, not believe red and white squares, so the sampling mode reads each shard’s 16 byte index entry, fetches the inner 32 by 32 chunk, runs it through the blosc codec chain, dequantises the first three bands and applies the store’s own RGB stretch. That is what the collage at the top of this post is made of.
Clearly, SUCCEEDED is not the same as complete. Job status tells you a process exited zero, not that it wrote what it was supposed to write!

The same 2025 map as it looked on the 24th. Each red stripe is a UTM zone that reported success. The narrower stripes through Africa and Asia stop dead at the equator, which is the signature that made the fault findable.
Pool capacity #
I mentioned last week about the three levels of queues with increasing capacity. Since all the remaining utm zones were the dense ones there was now a backlog on the top d16v tier with the lower tiers empty. The d16v pool was raised from 128 vCPU to 256, then 400, then 640.
Three zone years ran out of memory even at 120 GB. The fix was to run these with a smaller read window.
AMX on Endeavour #
This section is a work in progress.
Intel gave us time on Endeavour so we could run Tessera inference on AMX and see how CPU compares with GPU.
The machines have no outbound internet access and no reverse tunnels so everything has to be staged.
Running the Manchester test tile, which I’d used on many previous occasions, grid_-1.65_53.35 for 2024, 1125 by 687 pixels, across every CPU type available with the node filled:
| CPU | tiles/h/node |
|---|---|
| Xeon 6980P, 2 by 128c | 56 |
| 6972P, 2 by 96c, MCR DIMMs | 51 |
| 6972P, 2 by 96c, DDR | 47 |
| 6966P-C, 2 by 96c | 46 |
| 6960P, 2 by 72c | 34 |
| 6767P, 2 by 64c | 31 |
| 6787P, 2 by 86c | 30 |
| Platinum 8592+, 2 by 64c | 25 |
| Platinum 8480+, 2 by 56c | 20 |
| 6745P, 2 by 32c | 20 |
| 6990E+, 2 by 288 E-cores, no AMX | 18 |
Two things stand out. bf16 on AMX is about 3.4 times faster than f32, and the part with 576 E-cores and no AMX at all comes last despite having by far the most cores. AMX is doing the work.
Uncontended, a single tile takes between 6.5 and 8 minutes on the AMX parts, against about 30 seconds on an MI355X. All the AMX parts land within about 10% of each other at eight threads, so per core they are the same speed and the node types differ only in how many cores they can feed.
day10 #
Base image system dependencies #
On Obuilder, we rebuild the base images every week, which ensures that their apt (or other package manager) cache has the latest mirrors available. Once the base image has been updated all cache layers are now obsolete and everything is rebuilt again. This isn’t a significant overhead in Obuilder as any dependency set difference is a cache miss.
For day10, particularly on slow RISC-V machines, the cache is valuable, and the cold startup is expensive. Early on in the development process, I discovered that you can chroot to the base layer and run apt update (or equivalent) to update the lists. There is now a refresh-base sub command for day10 which does that via runc.
ocluster-worker, which already monitors disk space and other metrics, now has a timer which uses the same pause and drain mechanism that pruning uses, to pause the worker and update the cache every 24 hours.
Depexts belong in the layer key #
One of the early ideas which allowed day10 to work was that the layer hash is built from the hash of the opam files it contains. opam already has a mechanism to provide the hash of an opam file called effective_part, which ignores insignificant changes.
Interestingly, opam does not consider depexts to be significant. This is absolutely correct because system packages do not change the artefact installed into a switch. It is wrong for day10, where installing the depexts is part of building the layer.
depexts = empty.depexts; (* opamFile.ml:3514 *)
The consequence showed up on ocaml/opam-repository#30785, which adds gmp-static to conf-gmp.5 and changes nothing else. The key did not move, a week old layer answered from the cache, and we got a false pass for conf-gmp and a false failure for the revdep the PR existed to fix.
day10 now builds its own effective_part up from empty with explicit setters, rather than editing opam’s value. That way a field opam adds later stays out until someone decides it belongs, and an opam upgrade cannot silently re-key every cache on every builder. That is not hypothetical as opam 2.5 changes effective_part.
z3 crashed doris #
On Wednesday morning, z3.5.1.0 OOMed the worker, and the machine was unreachable. It was not one greedy process but the aggregate of many cc1plus processes.
Eighteen distributions, each with two compilers, yields 36 z3 jobs in the matrix. The worker has a capacity of 64 concurrent jobs. opam gives each build one job per core, and a container sees all 256 of doris’s cores, so each of those builds runs make -j 255.
Neither day10 nor OBuilder caps this, so it is not a regression. It is specific to doris because OBuilder spreads those 36 variants across many workers, whereas the day10 shadow puts the whole amd64 matrix on one box.
Implementing a fixed OPAMJOBS would be wrong, because a four-core board would be crippled by any number which makes doris happy. Deriving the number from memory rather than cores, 750 GB divided by 64 concurrent builds at 375 MB gives about 31, so the cap is 32:
OPAMJOBS = max 1 (min (cores - 1) 32)
carpenter and the RISC-V boards come out at 3 as before and are untouched. Only doris changed. Afterwards, the same workload ran 576 concurrent cc1plus processes using 138 GB of 1007 GB, with 869 GB free.
This does have a performance penalty: testing with an idle machine, it takes 68 seconds to build z3 at -j 255 against 126 at -j 32. This is a significant difference, but the difference won’t typically be that large as it’s unusual for a build to be able to compile that many individual files at the same time.
The platform day10 invented #
All of the day10 --os... parameters are derived from the machine if not specified. Neither opam-repo-ci nor ocaml-ci passed --os-family as it was not something which was readily available, so every job got the default.
That was harmless while depexts were not in the layer key, because inside the container opam reads the container’s own /etc/os-release and installs the right packages regardless. Once depexts became part of the key, it decided the key. conf-clang.2 on alpine resolved to both clang21 and clang, the second from a debian family line that only matched because of the guess, and the same package on the same platform keyed differently depending on whether the flag was passed:
--os-distribution alpine --os-version 3.24 -> 46618c93...
--os-distribution alpine --os-version 3.24 --os-family alpine -> c3297387...
When --os-family isn’t specified it is now derived from --os and --os-distribution in commit ef18a3b34b43ae25cb792af3f2096504ccf6b2c1.
A green tick for tests that never ran #
A runtest job is distinguishable from a build job of the same package because the test flag is in the layer key. The problem was that day10 only flagged a layer as a test layer when some of the command lines included the {with-test} filter.
Layers which didn’t have that and ran tests due to run-test: [ "dune" "runtest" ] replayed the cache of the build layer without running any tests.
Of 19,113 opam files, 549 use run-test and of those 188 don’t mention with-test anywhere. Those 188 would replay the build layer without running the test. run-test is now an additional marker along with {with-test} closing this gap.
PID 1 #
ppx_windtrap.0.2.0 failed on day10 on every distribution and architecture, and passed on OBuilder everywhere. Three tests in the same file, all hitting a ten second timeout:
FAIL Limits > an expired child's process group dies whole
expected ["gone"] actual ["alive"]
The namespaces were byte-identical to OBuilder’s. The difference is what runs as PID 1. day10 runs the build as a single shell command, and bash execs a lone simple command rather than forking it, so the package’s own build process became PID 1. PID 1 is the only thing that reaps orphans, and a zombie still answers kill(pid, 0), so a test suite that kills a process group and waits for it to disappear sees it stay alive forever.
OBuilder passes by accident, because its script has more than one command in it, which leaves a shell as PID 1.
The fix is to make a trivial change which stops bash seeing it as a simple command:
- let argv = [ "/usr/bin/env"; "bash"; "-c"; command ] in
+ let argv = [ "/usr/bin/env"; "bash"; "-c"; command ^ "\nexit $?" ] in
The suite now passes and takes 3.6 seconds rather than 53.7. I am not entirely happy with the fix. A real init in the container is the proper answer. tini is packaged on only 14 of our 19 platforms and has not had a release since 2020, catatonit covers 18, so the likely route is building it as a multi-stage and copying the binary in, using machinery the Dockerfile already has.
Two builds of the same package are not the same build #
conf-libclang.22 builds with bash -ex configure.sh version, where version is an opam variable. Installing the package writes a .config that republishes version as the detected LLVM version, so the second build resolves a different command:
+ bash "-ex" "configure.sh" "22" <- install
+ bash "-ex" "configure.sh" "19.1.7" <- test run, fails
And because that package has no run-test field, the build job and the test job share one layer, so whichever ran first decided the outcome for both.
The fix restores the invariant rather than dodging it, by having day10-install reinstall the package for the test run, which is what opam reinstall does and for the same reason.
The same thread turned up a depext filter bug. day10 counted a filter it could not evaluate as applying, on the grounds that over-keying costs only a rebuild. But opam decides what actually gets installed, and it counts the same filter as not applying, so the key described packages that were never installed. A package variable in a depexts filter parses but never resolves, which is ocaml/opam#5836; goblint carries such a line with a comment saying it does not work, and about forty more reach for npm-version.
Containers no longer share the network #
Every container shared the machine’s network namespace. On a builder running dozens of jobs at once, two test suites binding the same fixed port collide, which retrospectively explains some unexplained EADDRINUSE in earlier logs. Worse, one job could reach another job’s server and report on what it found there. A suite that downloads something also passed under day10 while failing under OBuilder, which takes the network away for exactly this reason.
So when tests are requested, the package is installed in one container, and the tests run in a second with no network. The second does not need one as the first installed the depexts and left the sources in the download cache, and both write into the same upper directory, so the build tree is still there.
The fix that was backwards #
openSUSE Tumbleweed failed conf-libtool with "libtoolize": command not found. This is not the CodeReady Builder situation from CentOS; Tumbleweed’s only disabled repos are debug and source, and libtool is in an enabled one. The difference was the base image: ocaml-dockerfile installs a development pattern or group, and day10’s image installs an explicit list, which does not include libtool.
So I made day10’s base images match OBuilder’s package sets, and then rejected my own change before it went anywhere. A base image fat enough to cover a wrong depext stops day10 finding the next one, which is backwards for a tool whose job is to surface exactly this. The underlying bug is sharper than “missing depext”, too:
centos:stream9 ID="centos" ID_LIKE="rhel fedora" -> opam os-family is rhel
conf-automake: [ "automake" ] {os-family = "centos"} -> never matches
centos is a distribution, not a family, so that line has never fired and automake was simply never installed on CentOS. conf-libtool had no SUSE entry at all. Both are one-line changes, and they went to opam-repository as #30792, surfaced by libdash.0.5.2 failing on four variants.
opam-build name confusion #
day10-install is the new name for the in-container installer, which I’d previously called opam-build, after Kate pointed out the collision with her own opam-build opam plugin. mtelvers/day10-install
CentOS CRB #
CentOS ships CodeReady Builder disabled, and zlib-static and zlib-ng-compat-static both live in it. ocurrent/ocaml-dockerfile enables it; day10’s generator never did. jmid reported this.
New prune option #
prune now accepts --failed, replacing a shell one-liner I kept running by hand.
Scheduler pools #
Up to now, I had been using spare pools on the scheduler for the day10 work but I had run out of unused pools and wanted to add ppc64 and arm64. Adding pools requires stopping the scheduler which cancels all in-progress jobs, so I try to avoid it whenever I can.
I was a little surprised that the pool name forms part of the cache key so updating the pool names from test to day10-linux-riscv64 actually caused all the jobs to be submitted again. Fortunately with a warm cache this wasn’t too bad.
Note that obuilder can coexist with day10 in the same worker, meaning that ocluster-worker could accept day10, obuilder and docker jobs, but I have chosen not to deploy it like that at the moment as it’s easier to debug and update when there are fewer machines involved.
ARM64 and POWER9 #
The day10 shadow now covers four architectures. The distribution set is not the same on each, because it is whatever ocaml-dockerfile publishes base images for: there are no ppc64le images for the Fedora, Alpine, openSUSE, Arch or CentOS families, so POWER9 is Debian and Ubuntu only, and RISC-V stays on debian-13 alone.
| Architecture | Distributions |
|---|---|
| x86_64 | 18 |
| ARM64 | 12 |
| POWER9 | 7 |
| RISC-V | 1 |
Each is built against both default compilers, 4.14 and 5.5.
The display tree in opam-repo-ci has been restructured to better show the distributions and platforms.
bench.ci.dev out of space #
The Docker disk was 100% full, and Postgres could not write its lock file. This is not the first time this has happened. docker system prune runs on a cron job so I knew that wasn’t going to save any space. In the past, I’ve always ended up deleting overlayfs and redeploying the containers.
This time, I decided to investigate in more depth. Every ocurrent/current-bench job leaks the image it builds: the worker builds with --iidfile and no tag, so the image is untagged from creation; containers are cleaned up because it runs with --rm; and the finaliser only unlinks the ID file. There is no docker rmi anywhere in it. The result was about 1,375 orphaned overlay directories holding roughly 140 GB, which docker system prune reports as nothing reclaimable because they are layerdb records with no image to anchor them. Delicate surgery took the disk from 183 GB used to 23 GB. The fix is eighteen lines and is PR #502.
FreeBSD / rosemary #
It’s been a few months (weeks?) since I’ve written about my FreeBSD woes. We were back to the situation where trivial FreeBSD jobs were timing out at the two-hour limit while the machine sat at load 0.06. As in previous investigations the culprit looked to be the filesystem, either ZFS itself or in FreeBSD vnodes.
Unmounting each result after it is built took resident vnodes from 1.5 million to 362,000 and became ocurrent/obuilder#220.
macOS, which also uses ZFS, had been attacking the same problem from the other end. Its sandbox’s finished () unmounted and remounted the whole obuilder/result dataset once at the end of a run, clearing every mount in one go rather than never letting them accumulate:
let finished () =
Os.sudo [ "zfs"; "unmount"; "obuilder/result" ] >>= fun () ->
Os.sudo [ "zfs"; "mount"; "obuilder/result" ] >>= fun () ->
Lwt.return ()
Unmounting each result as it is snapshotted makes that redundant, so it is removed and falls back to the default no-op. That also drops a sandbox reaching into the store, which it had no business doing.
The evidence that it is the right change is narrow. I tried at length to reproduce the actual problem, but the closest I got was a catastrophic slowdown under free vnode exhaustion.
In order to fix some of the corrupt obuilder cache layers, I needed to wipe the obuilder store, which even without any fix would get us several weeks of good behaviour.
I added a cron job to report on various metrics on ZFS and vnodes to see if I can see any trends in the data. I also wanted to have some measure of how the file system was performing on normal git operations. I considered writing a test zfs clone / git clone script, but concluded that I didn’t need one. Obuilder jobs which have no solution on FreeBSD, do exactly what I want, clone with ZFS file system, update the git clone and then exit. I set up a cron job to pull these build times from opam-repo-ci’s log.
Five days in is not long, but so far everything is as normal.
Windows: ltsc2019 onto the HCS store #
ltsc2019-1 failed twice in the window with near identical symptoms and completely different causes, which is worth separating.
On Monday, it had livelocked: 20,059 prune cycles and 40,096 Pruned 0 items, obuilder’s store holding a single item so its prune could never free anything, while Docker sat on 72.77 GB it would not release until its own much lower threshold. Pruning Docker recovered 71 GB, but the restart then failed its self-test with UNIQUE constraint failed: builds.id, a database row whose result the prune had deleted, with an uncheckpointed write-ahead log stamped at the moment it died mid livelock.
I should add here that ltsc2019-1 was running Obuilder backed by Docker layers and had never been migrated to HCS. I hadn’t migrated it because I was sure containerd or something in HCS would be different between LTSC 2025 and 2019 which would need a lot of detailed investigation. Obuilder backed by Docker layers has some issues monitoring free space, as both Obuilder spec jobs and Docker jobs use the same Docker store, and free space triggers a global docker system prune, which obviously affects both stores.
Since I had Ansible playbooks for the entire Windows build process, I decided to try HCS on LTSC 2019 and see where the failures were. I built ltsc2019-2 alongside the existing machine, and much to my surprise, the HCS implementation worked without any changes.
NFS out of space #
daintree’s ZFS backed NFS volume /maps reached 100% full. Users deleted files, but because of the ZFS snapshots nothing actually frees. Back in April, I had planned to move /maps to Ceph and did an initial copy of all the data. With writes now blocked by space, I refreshed my copy of the 127 million files!
I also applied some tuning to the Ceph deployment, increasing to two active MDS both with standby-replay.
Prometheus 2 and ocurrent #
Applying some minor code tidy-ups in ocurrent, I noticed that the CI failed everywhere due to the changes in Prometheus 2. Accepting that I really should be moving to Prometheus 2, I opted for the easy route of applying an upper bound just to get my other PR to clear.
Thomas, talex5, pointed out that if nothing reports metrics through the Lwt interface, then the synchronous collector is all that is needed, so the upgrade path was actually straightforward.
Using day10, I could easily simulate a world where ocurrent used Prometheus 2.0 and see what other builds failed, specifically looking at opam-repo-ci and OCaml-CI. opam-repo-ci was fine, but OCaml-CI failed due to the solver-service. The PR there looks straightforward ocurrent/solver-service#84
- let open Lwt.Infix in
let content_type = "text/plain; version=0.0.4; charset=utf-8" in
- Prometheus.CollectorRegistry.(collect default) >>= fun data ->
+ let data = Prometheus.CollectorRegistry.(collect default) in
Therefore, I will now move ocurrent to Prometheus 2 in ocurrent/ocurrent/pull/479.
32-bit OCaml and Intel macOS #
The 32-bit compiler fork’s nightly rebase had been green for twenty consecutive days and then failed the rebase. Upstream has updated the README, moving macOS on x86_64 from Tier 1 to Tier 2 in the platform table. Resolved by hand, 17 commits replayed cleanly. Intel Macs shipped from 2006 to 2020, and Homebrew have moved them to the second tier too. It is the end of an era.
Minor tidy up PRs #
I filed a few tidy-up PRs to harden some of our software: ocurrent/ocurrent/pull/477, ocurrent/opam-repo-ci/pull/486, ocurrent/opam-repo-ci/pull/487, ocurrent/obuilder/pull/221 and ocurrent/obuilder/pull/222.
occ on RISC-V #
The RISC-V back end got some additional attention. The assembler now covers RV64, including the address pseudo-instructions, assembler macros, the fused multiply-adds and upper case mnemonics. The linker produces static executables, dynamic executables and shared objects, and the version line no longer says the linker cannot do this machine.
The OCaml testsuite goes 1,617 passed, 61 skipped, 1 failed, against a gcc baseline of 1,621 passed, 58 skipped, and 0 failed.
The one remaining RISC-V failure is lazy3.ml in bytecode, which does not produce an incorrect answer but times out. The interpreter measured 6.2 times slower than gcc’s, which scaled up to about 595 seconds against a 600-second limit.