mirror of
https://github.com/apache/nuttx.git
synced 2026-10-11 16:20:21 +00:00
Rewrite arch_memcpy.S and arch_memset.S to be register-width aware on both RV32 and RV64 using REG_L/REG_S/SZREG macros from asm.h. memcpy gains: - 16xSZREG unrolled main loop (128B/iter on RV64, 64B on RV32). - Shift-merge path for misaligned src: reads two aligned words straddling each output word and shifts them together, so no load or store is ever misaligned. - Single SZREG and byte loops for remainder and small copies. memset gains: - 32xSZREG unrolled main loop (256B/iter on RV64, 128B on RV32) using Duff's device for non-power-of-two remainders. - .option norvc ensures fixed 4-byte instruction width for correct jump offset calculation in the Duff's device entry. - Zero-length input handled correctly (branch to guarded tail). The old memcpy always used lw/sw even on RV64, wasting half the memory bandwidth. The old memset unrolled only 16 bytes per iteration. Signed-off-by: ganjing <ganjing@xiaomi.com> |
||
|---|---|---|
| .. | ||
| arch_elf.c | ||
| arch_memchr.S | ||
| arch_memcmp.S | ||
| arch_memcpy.S | ||
| arch_memmove.S | ||
| arch_memset.S | ||
| arch_setjmp.S | ||
| arch_stpcpy.S | ||
| arch_stpncpy.S | ||
| arch_strcat.S | ||
| arch_strchr.S | ||
| arch_strchrnul.S | ||
| arch_strcmp.S | ||
| arch_strcpy.S | ||
| arch_strlcpy.S | ||
| arch_strlen.S | ||
| arch_strncmp.S | ||
| arch_strncpy.S | ||
| arch_strnlen.S | ||
| arch_strrchr.S | ||
| asm.h | ||
| CMakeLists.txt | ||
| Kconfig | ||
| Make.defs | ||