Ladybug version
main
What operating system are you using?
ubuntu24.04
What happened?
Conclusion: This was not an OOM event. It was a SIGSEGV (segmentation fault) triggered by the LadybugDB native. The most likely cause is a concurrency or lifecycle race between repeatedly reopening the read-only database, closing old connection pools, and concurrent checkpoint/WAL operations.
Key evidence:
-
The log explicitly reports:
Fatal Python error: Segmentation fault
-
The crashing thread was in:
ladybug/database.py -> init_pybind_database
Call chain:
follower_refresh_loop -> _refresh_follower_main_database -> _open_main_database -> Ladybug Database initialization
This indicates that the crash occurred in Ladybug’s native/pybind layer, rather than in normal Python application code.
-
The database was repeatedly reconnected, approximately once every 15 seconds. The read_pool_generation increased from 1 to 19. Each refresh performed the following operations:
- Opened the same
main_ontology.lbug database.
- Created 8 read connections.
- Switched the active connection pool.
- Closed the previous database instance.
-
Before the crash, the log showed database-file concurrency/state errors:
Cannot open database in read-only mode while checkpoint is in progress
Cannot open file ... main_ontology.lbug.wal: No such file or directory
These errors indicate that the database was being opened for reading while its files may have been changing due to a checkpoint, WAL removal/rotation, or another process modifying the database.
-
Memory usage did not reach the container limit:
- cgroup memory limit: 16 GiB
- Last recorded cgroup usage: approximately 2.31 GiB
- Process RSS: approximately 1.97 GiB
- Memory monitor status:
normal
Therefore, this does not resemble a Kubernetes OOM kill. An OOM kill normally results in OOMKilled and exit code 137, while a segmentation fault commonly results in exit code 139.
Most likely failure sequence:
Checkpoint/WAL state changes
↓
Read-only database opening fails and triggers retries
↓
Repeated open + connection-pool switch + close operations every 15 seconds
↓
A thread-safety or resource-lifecycle issue occurs in the Ladybug native layer
↓
SIGSEGV / core dump
Based on the current logs, it can be confirmed that the segmentation fault occurred during Ladybug native database initialization. The strongest suspected root cause is frequent read-database reconnection while checkpoint/WAL operations are changing the same database files, causing a thread-safety or resource-release issue in the native layer.
Are there known steps to reproduce?
I’m not sure how to go about fixing this issue. Every time I observe the problem, I only get a static core dump, and I cannot reproduce the failure reliably. As a result, it is unclear where exactly the root cause lies. Could you share some suggestions and troubleshooting ideas? @adsharma
Ladybug version
main
What operating system are you using?
ubuntu24.04
What happened?
Conclusion: This was not an OOM event. It was a
SIGSEGV(segmentation fault) triggered by the LadybugDB native. The most likely cause is a concurrency or lifecycle race between repeatedly reopening the read-only database, closing old connection pools, and concurrent checkpoint/WAL operations.Key evidence:
The log explicitly reports:
Fatal Python error: Segmentation faultThe crashing thread was in:
ladybug/database.py -> init_pybind_databaseCall chain:
follower_refresh_loop -> _refresh_follower_main_database -> _open_main_database -> Ladybug Database initializationThis indicates that the crash occurred in Ladybug’s native/pybind layer, rather than in normal Python application code.
The database was repeatedly reconnected, approximately once every 15 seconds. The
read_pool_generationincreased from 1 to 19. Each refresh performed the following operations:main_ontology.lbugdatabase.Before the crash, the log showed database-file concurrency/state errors:
Cannot open database in read-only mode while checkpoint is in progressCannot open file ... main_ontology.lbug.wal: No such file or directoryThese errors indicate that the database was being opened for reading while its files may have been changing due to a checkpoint, WAL removal/rotation, or another process modifying the database.
Memory usage did not reach the container limit:
normalTherefore, this does not resemble a Kubernetes OOM kill. An OOM kill normally results in
OOMKilledand exit code 137, while a segmentation fault commonly results in exit code 139.Most likely failure sequence:
Based on the current logs, it can be confirmed that the segmentation fault occurred during Ladybug native database initialization. The strongest suspected root cause is frequent read-database reconnection while checkpoint/WAL operations are changing the same database files, causing a thread-safety or resource-release issue in the native layer.
Are there known steps to reproduce?
I’m not sure how to go about fixing this issue. Every time I observe the problem, I only get a static core dump, and I cannot reproduce the failure reliably. As a result, it is unclear where exactly the root cause lies. Could you share some suggestions and troubleshooting ideas? @adsharma