earlephilhower
(Migrated from github.com)
left a comment
Copy Link
Copy Source
Thanks, but this is pretty expensive in terms of ROM vs. what it's (hopefully) accelerating. 256 bytes of bytewide 1s and 0s just feels wasteful. And since it's in ROM it could be a very slow access (XIP is nice, but nobody's going to call it fast...).
I'm not sure that parity on any bitrate that the SerialPIO operates at is a limiting factor, but if you really think it is then why not something simpler with much lower space requirements.
static int __not_in_flash_func(_parity)(int data) {
data ^= data >> 4;
data &= 0xf;
return (0x6996 >> data) & 1;
}
There shouldn't be a need to mask off bits before doing the math in the above or a LUT. That was there to short-circuit in the original for loop case.
Thanks, but this is pretty expensive in terms of ROM vs. what it's (hopefully) accelerating. 256 bytes of bytewide 1s and 0s just feels wasteful. And since it's in ROM it could be a very slow access (XIP is nice, but nobody's going to call it fast...).
I'm not sure that parity on any bitrate that the SerialPIO operates at is a limiting factor, but if you really think it is then why not something simpler with much lower space requirements.
For example, (stolen from http://www.graphics.stanford.edu/%7Eseander/bithacks.html#ParityParallel) you can do parity in ~5 instructions without LUT by utilizing a constant as a parity bitmap:
````
static int __not_in_flash_func(_parity)(int data) {
data ^= data >> 4;
data &= 0xf;
return (0x6996 >> data) & 1;
}
````
There shouldn't be a need to mask off bits before doing the math in the above or a LUT. That was there to short-circuit in the original for loop case.
that bit hack is way more brilliant. I'm using this for running 8 rx channels on my rp2040 with real-time constraints alongside other critical tasks. cpu usage matters when you're maxing out a cortex-m0 cpu because that cpu has potential indeed
that bit hack is way more brilliant. I'm using this for running 8 rx channels on my rp2040 with real-time constraints alongside other critical tasks. cpu usage matters when you're maxing out a cortex-m0 cpu because that cpu has potential indeed
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
use LUT for faster parity checks, reduces CPU load with multiple instances
Thanks, but this is pretty expensive in terms of ROM vs. what it's (hopefully) accelerating. 256 bytes of bytewide 1s and 0s just feels wasteful. And since it's in ROM it could be a very slow access (XIP is nice, but nobody's going to call it fast...).
I'm not sure that parity on any bitrate that the SerialPIO operates at is a limiting factor, but if you really think it is then why not something simpler with much lower space requirements.
For example, (stolen from http://www.graphics.stanford.edu/%7Eseander/bithacks.html#ParityParallel) you can do parity in ~5 instructions without LUT by utilizing a constant as a parity bitmap:
There shouldn't be a need to mask off bits before doing the math in the above or a LUT. That was there to short-circuit in the original for loop case.
that bit hack is way more brilliant. I'm using this for running 8 rx channels on my rp2040 with real-time constraints alongside other critical tasks. cpu usage matters when you're maxing out a cortex-m0 cpu because that cpu has potential indeed
Thx, LGTM!