Skip to content

Bug: Segmentation Fault During Periodic Read-Only Ladybug Database Reconnection #1007

Description

@ericyuanhui

Ladybug version

main

What operating system are you using?

ubuntu24.04

What happened?

Conclusion: This was not an OOM event. It was a SIGSEGV (segmentation fault) triggered by the LadybugDB native. The most likely cause is a concurrency or lifecycle race between repeatedly reopening the read-only database, closing old connection pools, and concurrent checkpoint/WAL operations.

Key evidence:

  • The log explicitly reports:

    Fatal Python error: Segmentation fault

  • The crashing thread was in:

    ladybug/database.py -> init_pybind_database

    Call chain:

    follower_refresh_loop -> _refresh_follower_main_database -> _open_main_database -> Ladybug Database initialization

    This indicates that the crash occurred in Ladybug’s native/pybind layer, rather than in normal Python application code.

  • The database was repeatedly reconnected, approximately once every 15 seconds. The read_pool_generation increased from 1 to 19. Each refresh performed the following operations:

    1. Opened the same main_ontology.lbug database.
    2. Created 8 read connections.
    3. Switched the active connection pool.
    4. Closed the previous database instance.
  • Before the crash, the log showed database-file concurrency/state errors:

    • Cannot open database in read-only mode while checkpoint is in progress
    • Cannot open file ... main_ontology.lbug.wal: No such file or directory

    These errors indicate that the database was being opened for reading while its files may have been changing due to a checkpoint, WAL removal/rotation, or another process modifying the database.

  • Memory usage did not reach the container limit:

    • cgroup memory limit: 16 GiB
    • Last recorded cgroup usage: approximately 2.31 GiB
    • Process RSS: approximately 1.97 GiB
    • Memory monitor status: normal

    Therefore, this does not resemble a Kubernetes OOM kill. An OOM kill normally results in OOMKilled and exit code 137, while a segmentation fault commonly results in exit code 139.

Most likely failure sequence:

Checkpoint/WAL state changes
        ↓
Read-only database opening fails and triggers retries
        ↓
Repeated open + connection-pool switch + close operations every 15 seconds
        ↓
A thread-safety or resource-lifecycle issue occurs in the Ladybug native layer
        ↓
SIGSEGV / core dump

Based on the current logs, it can be confirmed that the segmentation fault occurred during Ladybug native database initialization. The strongest suspected root cause is frequent read-database reconnection while checkpoint/WAL operations are changing the same database files, causing a thread-safety or resource-release issue in the native layer.

Are there known steps to reproduce?

I’m not sure how to go about fixing this issue. Every time I observe the problem, I only get a static core dump, and I cannot reproduce the failure reliably. As a result, it is unclear where exactly the root cause lies. Could you share some suggestions and troubleshooting ideas? @adsharma

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions