Skip to content

feat(riscv): Support advanced interrupt architecture - #2689

Closed
sbutz wants to merge 8 commits into
hermit-os:mainfrom
sbutz:sb/riscv_aia
Closed

sbutz wants to merge 8 commits into
hermit-os:mainfrom
sbutz:sb/riscv_aia

Conversation

@sbutz

@sbutz sbutz commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Implements #2614.

Feature flag is added in hermit-os/hermit-rs#1077.

Follow-up tasks:

  • For each configured msix vector. The external interrupt with the same number is enabled too. This might incur unnecessary cpu traps. Proposed solution: Implement msi vector allocator and store handlers per imsic.
  • Bump riscv crate once new version available.
  • Specify qemu cpu profile rva23s64 once qemu reaches version 10.1 on ci

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmark Results

Details
Benchmark Current: c9e016e Previous: 2e23902 Performance Ratio
startup_benchmark Build Time 92.20 s 80.34 s 1.15 ❗
startup_benchmark File Size 0.79 MB 0.80 MB 0.99 ❗
Startup Time - 1 core 0.74 s (±0.02 s) 0.75 s (±0.02 s) 0.99
Startup Time - 2 cores 0.76 s (±0.03 s) 0.74 s (±0.02 s) 1.04
Startup Time - 4 cores 0.75 s (±0.02 s) 0.74 s (±0.02 s) 1.01
multithreaded_benchmark Build Time 93.20 s 82.11 s 1.14 ❗
multithreaded_benchmark File Size 0.87 MB 0.86 MB 1.02 ❗
Multithreaded Pi Efficiency - 2 Threads 89.61 % (±8.97 %) 85.89 % (±6.61 %) 1.04
Multithreaded Pi Efficiency - 4 Threads 44.39 % (±3.10 %) 43.43 % (±2.56 %) 1.02
Multithreaded Pi Efficiency - 8 Threads 25.56 % (±1.23 %) 25.76 % (±1.53 %) 0.99
micro_benchmarks Build Time 92.65 s 80.40 s 1.15 ❗
micro_benchmarks File Size 0.88 MB 0.86 MB 1.02 ❗
Scheduling time - 1 thread 69.32 ticks (±3.97 ticks) 62.65 ticks (±4.06 ticks) 1.11 ❗
Scheduling time - 2 threads 37.97 ticks (±4.52 ticks) 34.08 ticks (±4.10 ticks) 1.11
Micro - Time for syscall (getpid) 3.89 ticks (±0.61 ticks) 3.45 ticks (±0.58 ticks) 1.13
Memcpy speed - (built_in) block size 4096 77590.11 MByte/s (±53651.66 MByte/s) 82448.38 MByte/s (±56997.13 MByte/s) 0.94
Memcpy speed - (built_in) block size 1048576 29527.85 MByte/s (±24016.75 MByte/s) 30585.98 MByte/s (±24707.84 MByte/s) 0.97
Memcpy speed - (built_in) block size 16777216 24698.97 MByte/s (±20507.99 MByte/s) 26340.06 MByte/s (±21720.96 MByte/s) 0.94
Memset speed - (built_in) block size 4096 78077.57 MByte/s (±53986.30 MByte/s) 82292.76 MByte/s (±56891.50 MByte/s) 0.95
Memset speed - (built_in) block size 1048576 30245.06 MByte/s (±24425.29 MByte/s) 31323.85 MByte/s (±25145.86 MByte/s) 0.97
Memset speed - (built_in) block size 16777216 25393.62 MByte/s (±20942.38 MByte/s) 27104.68 MByte/s (±22209.94 MByte/s) 0.94
Memcpy speed - (rust) block size 4096 69733.19 MByte/s (±48596.37 MByte/s) 74097.96 MByte/s (±51811.44 MByte/s) 0.94
Memcpy speed - (rust) block size 1048576 29454.31 MByte/s (±24010.79 MByte/s) 30361.60 MByte/s (±24602.37 MByte/s) 0.97
Memcpy speed - (rust) block size 16777216 24572.43 MByte/s (±20359.96 MByte/s) 27625.34 MByte/s (±22806.88 MByte/s) 0.89
Memset speed - (rust) block size 4096 70172.07 MByte/s (±48930.90 MByte/s) 74373.47 MByte/s (±51976.48 MByte/s) 0.94
Memset speed - (rust) block size 1048576 30240.74 MByte/s (±24472.45 MByte/s) 31110.89 MByte/s (±25033.24 MByte/s) 0.97
Memset speed - (rust) block size 16777216 25266.40 MByte/s (±20794.03 MByte/s) 28386.93 MByte/s (±23265.03 MByte/s) 0.89
alloc_benchmarks Build Time 86.71 s 74.76 s 1.16 ❗
alloc_benchmarks File Size 0.87 MB 0.87 MB 1.00 ❗
Allocations - Allocation success 91.38 % 91.31 % 1.00 ❗
Allocations - Deallocation success 100.00 % 100.00 % 1
Allocations - Pre-fail Allocations 61.60 % 61.44 % 1.00 ❗
Allocations - Average Allocation time 6250.68 Ticks (±379.89 Ticks) 5860.58 Ticks (±98.43 Ticks) 1.07
Allocations - Average Allocation time (no fail) 7215.59 Ticks (±395.18 Ticks) 6554.81 Ticks (±92.86 Ticks) 1.10 ❗
Allocations - Average Deallocation time 3070.82 Ticks (±564.94 Ticks) 1805.01 Ticks (±250.35 Ticks) 1.70 ❗
mutex_benchmark Build Time 113.38 s 79.82 s 1.42 ❗
mutex_benchmark File Size 0.88 MB 0.86 MB 1.02 ❗
Mutex Stress Test Average Time per Iteration - 1 Threads 13.02 ns (±0.37 ns) 12.10 ns (±0.41 ns) 1.08 ❗
Mutex Stress Test Average Time per Iteration - 2 Threads 89.14 ns (±3.03 ns) 40.26 ns (±1.68 ns) 2.21 ❗

This comment was automatically generated by workflow using github-action-benchmark.

@sbutz
sbutz force-pushed the sb/riscv_aia branch 3 times, most recently from e1eb0f8 to a129717 Compare September 1, 2026 14:59
@sbutz
sbutz marked this pull request as ready for review September 2, 2026 15:57
@sbutz
sbutz marked this pull request as draft September 11, 2026 10:58
use crate::scheduler::PerCoreSchedulerExt;

// Claim interrupt
let Some(irq) = EXTERNAL_INTERRUPT_CONTROLLER

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's one APLIC or PLIC in the system. Therefore locking the global interrupt controller was reasonable.
However if the system fully supports AIA, there's an IMSIC per core and therefore no need to lock the platform level interrupt controller to claim an interrupt.

let interrupt_file_addr = INTERRUPT_FILES.get().unwrap()[hart_id];
let mut interrupt_file =
unsafe { VolatileRef::new(NonNull::new(interrupt_file_addr.as_mut_ptr()).unwrap()) };
interrupt_file

@sbutz sbutz Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Missing fence before ipi.

When software sends IPIs by writing MSIs to the IMSICs of other harts, programmers should consider also the need to execute a FENCE instruction before each store instruction that writes such an MSI. In the absence of FENCEs, many systems guarantee to preserve the order of a hart’s loads and stores only to/from individual devices, not among multiple devices, and not at all for accesses to main memory. With such a system, it must be remembered that each IMSIC is likely to be considered a separate device among the many. For example, if hart A wants to notify hart B that it has completed a task involving accesses to some I/O device, hart A may need to execute a FENCE before sending an MSI to B’s IMSIC, to ensure that all of A’s accesses to the device have actually completed before the MSI could arrive at B. Similarly, if hart A stores anything to memory that should be visible at hart B, a FENCE is likely needed before a subsequent store sending an MSI to B’s IMSIC.

@sbutz

sbutz commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Closed in favor of smaller PRs: #2731

@sbutz sbutz closed this Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant