Repository navigation
Conversation
There was a problem hiding this comment.
Benchmark Results
Details
| Benchmark | Current: c9e016e | Previous: 2e23902 | Performance Ratio |
|---|---|---|---|
| startup_benchmark Build Time | 92.20 s |
80.34 s |
1.15 ❗ |
| startup_benchmark File Size | 0.79 MB |
0.80 MB |
0.99 ❗ |
| Startup Time - 1 core | 0.74 s (±0.02 s) |
0.75 s (±0.02 s) |
0.99 |
| Startup Time - 2 cores | 0.76 s (±0.03 s) |
0.74 s (±0.02 s) |
1.04 |
| Startup Time - 4 cores | 0.75 s (±0.02 s) |
0.74 s (±0.02 s) |
1.01 |
| multithreaded_benchmark Build Time | 93.20 s |
82.11 s |
1.14 ❗ |
| multithreaded_benchmark File Size | 0.87 MB |
0.86 MB |
1.02 ❗ |
| Multithreaded Pi Efficiency - 2 Threads | 89.61 % (±8.97 %) |
85.89 % (±6.61 %) |
1.04 |
| Multithreaded Pi Efficiency - 4 Threads | 44.39 % (±3.10 %) |
43.43 % (±2.56 %) |
1.02 |
| Multithreaded Pi Efficiency - 8 Threads | 25.56 % (±1.23 %) |
25.76 % (±1.53 %) |
0.99 |
| micro_benchmarks Build Time | 92.65 s |
80.40 s |
1.15 ❗ |
| micro_benchmarks File Size | 0.88 MB |
0.86 MB |
1.02 ❗ |
| Scheduling time - 1 thread | 69.32 ticks (±3.97 ticks) |
62.65 ticks (±4.06 ticks) |
1.11 ❗ |
| Scheduling time - 2 threads | 37.97 ticks (±4.52 ticks) |
34.08 ticks (±4.10 ticks) |
1.11 |
| Micro - Time for syscall (getpid) | 3.89 ticks (±0.61 ticks) |
3.45 ticks (±0.58 ticks) |
1.13 |
| Memcpy speed - (built_in) block size 4096 | 77590.11 MByte/s (±53651.66 MByte/s) |
82448.38 MByte/s (±56997.13 MByte/s) |
0.94 |
| Memcpy speed - (built_in) block size 1048576 | 29527.85 MByte/s (±24016.75 MByte/s) |
30585.98 MByte/s (±24707.84 MByte/s) |
0.97 |
| Memcpy speed - (built_in) block size 16777216 | 24698.97 MByte/s (±20507.99 MByte/s) |
26340.06 MByte/s (±21720.96 MByte/s) |
0.94 |
| Memset speed - (built_in) block size 4096 | 78077.57 MByte/s (±53986.30 MByte/s) |
82292.76 MByte/s (±56891.50 MByte/s) |
0.95 |
| Memset speed - (built_in) block size 1048576 | 30245.06 MByte/s (±24425.29 MByte/s) |
31323.85 MByte/s (±25145.86 MByte/s) |
0.97 |
| Memset speed - (built_in) block size 16777216 | 25393.62 MByte/s (±20942.38 MByte/s) |
27104.68 MByte/s (±22209.94 MByte/s) |
0.94 |
| Memcpy speed - (rust) block size 4096 | 69733.19 MByte/s (±48596.37 MByte/s) |
74097.96 MByte/s (±51811.44 MByte/s) |
0.94 |
| Memcpy speed - (rust) block size 1048576 | 29454.31 MByte/s (±24010.79 MByte/s) |
30361.60 MByte/s (±24602.37 MByte/s) |
0.97 |
| Memcpy speed - (rust) block size 16777216 | 24572.43 MByte/s (±20359.96 MByte/s) |
27625.34 MByte/s (±22806.88 MByte/s) |
0.89 |
| Memset speed - (rust) block size 4096 | 70172.07 MByte/s (±48930.90 MByte/s) |
74373.47 MByte/s (±51976.48 MByte/s) |
0.94 |
| Memset speed - (rust) block size 1048576 | 30240.74 MByte/s (±24472.45 MByte/s) |
31110.89 MByte/s (±25033.24 MByte/s) |
0.97 |
| Memset speed - (rust) block size 16777216 | 25266.40 MByte/s (±20794.03 MByte/s) |
28386.93 MByte/s (±23265.03 MByte/s) |
0.89 |
| alloc_benchmarks Build Time | 86.71 s |
74.76 s |
1.16 ❗ |
| alloc_benchmarks File Size | 0.87 MB |
0.87 MB |
1.00 ❗ |
| Allocations - Allocation success | 91.38 % |
91.31 % |
1.00 ❗ |
| Allocations - Deallocation success | 100.00 % |
100.00 % |
1 |
| Allocations - Pre-fail Allocations | 61.60 % |
61.44 % |
1.00 ❗ |
| Allocations - Average Allocation time | 6250.68 Ticks (±379.89 Ticks) |
5860.58 Ticks (±98.43 Ticks) |
1.07 |
| Allocations - Average Allocation time (no fail) | 7215.59 Ticks (±395.18 Ticks) |
6554.81 Ticks (±92.86 Ticks) |
1.10 ❗ |
| Allocations - Average Deallocation time | 3070.82 Ticks (±564.94 Ticks) |
1805.01 Ticks (±250.35 Ticks) |
1.70 ❗ |
| mutex_benchmark Build Time | 113.38 s |
79.82 s |
1.42 ❗ |
| mutex_benchmark File Size | 0.88 MB |
0.86 MB |
1.02 ❗ |
| Mutex Stress Test Average Time per Iteration - 1 Threads | 13.02 ns (±0.37 ns) |
12.10 ns (±0.41 ns) |
1.08 ❗ |
| Mutex Stress Test Average Time per Iteration - 2 Threads | 89.14 ns (±3.03 ns) |
40.26 ns (±1.68 ns) |
2.21 ❗ |
This comment was automatically generated by workflow using github-action-benchmark.
e1eb0f8 to
a129717
Compare
Get access to stopei::read_clear and sireg.
Send IPI via imsic.
| use crate::scheduler::PerCoreSchedulerExt; | ||
|
|
||
| // Claim interrupt | ||
| let Some(irq) = EXTERNAL_INTERRUPT_CONTROLLER |
There was a problem hiding this comment.
There's one APLIC or PLIC in the system. Therefore locking the global interrupt controller was reasonable.
However if the system fully supports AIA, there's an IMSIC per core and therefore no need to lock the platform level interrupt controller to claim an interrupt.
| let interrupt_file_addr = INTERRUPT_FILES.get().unwrap()[hart_id]; | ||
| let mut interrupt_file = | ||
| unsafe { VolatileRef::new(NonNull::new(interrupt_file_addr.as_mut_ptr()).unwrap()) }; | ||
| interrupt_file |
There was a problem hiding this comment.
Missing fence before ipi.
When software sends IPIs by writing MSIs to the IMSICs of other harts, programmers should consider also the need to execute a FENCE instruction before each store instruction that writes such an MSI. In the absence of FENCEs, many systems guarantee to preserve the order of a hart’s loads and stores only to/from individual devices, not among multiple devices, and not at all for accesses to main memory. With such a system, it must be remembered that each IMSIC is likely to be considered a separate device among the many. For example, if hart A wants to notify hart B that it has completed a task involving accesses to some I/O device, hart A may need to execute a FENCE before sending an MSI to B’s IMSIC, to ensure that all of A’s accesses to the device have actually completed before the MSI could arrive at B. Similarly, if hart A stores anything to memory that should be visible at hart B, a FENCE is likely needed before a subsequent store sending an MSI to B’s IMSIC.
|
Closed in favor of smaller PRs: #2731 |
Implements #2614.
Feature flag is added in hermit-os/hermit-rs#1077.
Follow-up tasks: