Add ARM assembly optimized memcpy for RP2350 #2552

Merged
earlephilhower merged 3 commits from memasm into master 2024-10-24 01:11:19 +03:00
earlephilhower commented 2024-10-22 05:37:38 +03:00 (Migrated from github.com)

33% faster for 4K memcpy using DMAMemcyp example

With this assembly:
CPU: 4835 clock cycles for 4K
DMA: 2169 clock cycles for 4K

Using stock Newlib memcpy:
CPU: 7314 clock cycles for 4K
DMA: 2175 clock cycles for 4K

(What's interesting is that if we place this in RAM it's actually slower in this test because the CPU instruction fetch will fight with the data read and write, causing stalls...neat!)

33% faster for 4K memcpy using DMAMemcyp example With this assembly: CPU: 4835 clock cycles for 4K DMA: 2169 clock cycles for 4K Using stock Newlib memcpy: CPU: 7314 clock cycles for 4K DMA: 2175 clock cycles for 4K (What's interesting is that if we place this in RAM it's actually slower in this test because the CPU instruction fetch will fight with the data read and write, causing stalls...neat!)
Sign in to join this conversation.