optimize parity calculations in SerialPIO (#2932) #2933

Merged
ahmedarif193 merged 1 commits from serial-pio-parity-update-1 into master 2025-05-02 00:06:52 +03:00
ahmedarif193 commented 2025-05-01 22:00:09 +03:00 (Migrated from github.com)

use LUT for faster parity checks, reduces CPU load with multiple instances

use LUT for faster parity checks, reduces CPU load with multiple instances
earlephilhower (Migrated from github.com) requested changes 2025-05-01 22:19:56 +03:00
earlephilhower (Migrated from github.com) left a comment

Thanks, but this is pretty expensive in terms of ROM vs. what it's (hopefully) accelerating. 256 bytes of bytewide 1s and 0s just feels wasteful. And since it's in ROM it could be a very slow access (XIP is nice, but nobody's going to call it fast...).

I'm not sure that parity on any bitrate that the SerialPIO operates at is a limiting factor, but if you really think it is then why not something simpler with much lower space requirements.

For example, (stolen from http://www.graphics.stanford.edu/%7Eseander/bithacks.html#ParityParallel) you can do parity in ~5 instructions without LUT by utilizing a constant as a parity bitmap:

static int __not_in_flash_func(_parity)(int data) {
    data ^= data >> 4;
    data &= 0xf;
    return (0x6996 >> data) & 1;
}

There shouldn't be a need to mask off bits before doing the math in the above or a LUT. That was there to short-circuit in the original for loop case.

Thanks, but this is pretty expensive in terms of ROM vs. what it's (hopefully) accelerating. 256 bytes of bytewide 1s and 0s just feels wasteful. And since it's in ROM it could be a very slow access (XIP is nice, but nobody's going to call it fast...). I'm not sure that parity on any bitrate that the SerialPIO operates at is a limiting factor, but if you really think it is then why not something simpler with much lower space requirements. For example, (stolen from http://www.graphics.stanford.edu/%7Eseander/bithacks.html#ParityParallel) you can do parity in ~5 instructions without LUT by utilizing a constant as a parity bitmap: ```` static int __not_in_flash_func(_parity)(int data) { data ^= data >> 4; data &= 0xf; return (0x6996 >> data) & 1; } ```` There shouldn't be a need to mask off bits before doing the math in the above or a LUT. That was there to short-circuit in the original for loop case.
ahmedarif193 commented 2025-05-01 23:34:33 +03:00 (Migrated from github.com)

that bit hack is way more brilliant. I'm using this for running 8 rx channels on my rp2040 with real-time constraints alongside other critical tasks. cpu usage matters when you're maxing out a cortex-m0 cpu because that cpu has potential indeed

that bit hack is way more brilliant. I'm using this for running 8 rx channels on my rp2040 with real-time constraints alongside other critical tasks. cpu usage matters when you're maxing out a cortex-m0 cpu because that cpu has potential indeed
earlephilhower (Migrated from github.com) approved these changes 2025-05-01 23:50:12 +03:00
earlephilhower (Migrated from github.com) left a comment

Thx, LGTM!

Thx, LGTM!
Sign in to join this conversation.