mirror of
https://github.com/apache/nuttx.git
synced 2026-08-14 00:43:13 +00:00
Rewrite arch_memcpy.S and arch_memset.S to be register-width aware on both RV32 and RV64 using REG_L/REG_S/SZREG macros from asm.h. memcpy gains: - 16xSZREG unrolled main loop (128B/iter on RV64, 64B on RV32). - Shift-merge path for misaligned src: reads two aligned words straddling each output word and shifts them together, so no load or store is ever misaligned. - Single SZREG and byte loops for remainder and small copies. memset gains: - 32xSZREG unrolled main loop (256B/iter on RV64, 128B on RV32) using Duff's device for non-power-of-two remainders. - .option norvc ensures fixed 4-byte instruction width for correct jump offset calculation in the Duff's device entry. - Zero-length input handled correctly (branch to guarded tail). The old memcpy always used lw/sw even on RV64, wasting half the memory bandwidth. The old memset unrolled only 16 bytes per iteration. Signed-off-by: ganjing <ganjing@xiaomi.com> |
||
|---|---|---|
| .. | ||
| arm | ||
| arm64 | ||
| mips | ||
| renesas | ||
| risc-v | ||
| sim | ||
| sparc | ||
| tricore | ||
| x86 | ||
| x86_64 | ||
| xtensa | ||
| arch_atomic.c | ||
| arch_libc.c | ||
| CMakeLists.txt | ||
| Kconfig | ||
| Make.defs | ||