Skip to content

Bug: published extension binaries are stale since 2026-09-01, so the #940 HNSW fix has never shipped #1010

Description

@blarghmatey

Ladybug version

0.20.4 from PyPI, and the vdev channel as of today.

What operating system are you using?

Ubuntu on WSL2, x86_64. The staleness is server side, so it is not specific to a platform.

What happened?

extension.ladybugdb.com is serving extension binaries built on 2026-09-01 for every channel, including vdev. LadybugDB/extensions#83, which fixed the HNSW visited-bitmap overflow behind #940, merged on 2026-09-09, so that fix has never reached a client through INSTALL.

I found this chasing heap corruption in our own ingest, which turned out to be #940 reached by a different trigger. We were on 0.20.4, released the day after #83 merged, and the crash still reproduced, which is what sent me looking at the binaries rather than at the source.

The origin is stale, not just the CDN edge:

$ curl -sI https://extension.ladybugdb.com/vdev/linux_amd64/vector/libvector.lbug_extension
HTTP/2 200
content-length: 1164136
last-modified: Tue, 01 Sep 2026 23:23:02 GMT
age: 2396

age: 2396 means that response was fetched from origin forty minutes earlier, and it still reports a September 1 mtime. The directory index agrees: every channel from v0.11.3 through vdev is dated 01-Sep-2026, 23:20 to 23:23.

The published image is current, so this looks like the deploy step rather than the build. ghcr.io/ladybugdb/extension-repo:latest reports created: 2026-09-21T17:57:42Z, and build-extension / deploy-extensions succeeded in the 2026-09-21 nightly. So either the container behind extension.ladybugdb.com has not been replaced since September 1, or the releases tree baked into the image is itself stale. I could not tell which without pulling the 2.4 GB layer.

The served binary is pre-#83 independently of any timestamp. #83 changed initQueryHNSWSharedState to size the visited bitmap from NodeTable::getNumTotalRows() rather than the estimated cardinality. Both that function and NodeTable::getStats are exported dynamic symbols in the engine, so neither inlines across the extension boundary:

$ nm -D --undefined-only ~/.lbdb/extension/0.20.0/linux_amd64/vector/libvector.lbug_extension \
    | c++filt | grep NodeTable::get
   U lbug::storage::NodeTable::getStats(lbug::transaction::Transaction const*) const

getNumTotalRows is absent. A post-#83 build would have to import it.

There is a second problem that will outlive the first. nginx sets expires 15552000s, so every object carries a 180 day max-age and Cloudflare honours it:

$ curl -sI https://extension.ladybugdb.com/v0.20.0/linux_amd64/vector/libvector.lbug_extension
last-modified: Tue, 01 Sep 2026 23:22:38 GMT
expires: Wed, 03 Mar 2027 10:42:23 GMT
cache-control: public, max-age=15552000, no-transform
age: 1569336

That object has been held at this PoP for 18 days and is not due to revalidate until March 2027. Redeploying the origin will not help a client on a PoP that already has the stale copy, so this needs a purge as well. There is a purge-extension.yml workflow in this repo, so the mechanism exists.

Expected behavior: INSTALL vector on 0.20.4, or on a nightly, gets an extension built from the submodule the engine pins, which on main is 30bc1f15 and has #83 as an ancestor.

The impact for us is that every vector-index query runs against the pre-fix bitmap, and our ingest aborts with glibc heap corruption roughly once per episode. We have worked around it by preparing each index-call query fresh instead of letting the client re-execute its cached statement, which costs a parse and depends on a prepare-then-execute API that ladybug-python deprecates, so it is not something we can sit on indefinitely.

Are there known steps to reproduce?

Two calls, no build required:

curl -sI https://extension.ladybugdb.com/vdev/linux_amd64/vector/libvector.lbug_extension
curl -s  https://extension.ladybugdb.com/ | grep -E 'v0\.20\.0|vdev'

Both report 01-Sep-2026.

For the crash this strands, on 0.20.4 against the currently served vector extension: seed a node table with 300 rows carrying a 768 dimension FLOAT[] property and an HNSW index, then loop a CALL QUERY_VECTOR_INDEX(...) query through conn.execute(query_string, params) with an insert into the indexed column between iterations. It aborts around iteration 15 with "corrupted size vs. prev_size" or a segfault, where re-preparing the statement each time survives 3,000. Happy to attach that script, though I suspect the stale binary is the whole story.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions