Please see #12 for more details. Tl;dr it seems like we can get a ~6.5x speedup on inference (at least on an L40S with the case of a 32x32 grid with 6 timesteps, one channel and batch size 4) by changing one line of code.
I'm using walrus for a personal project so unfortunately I don't have much time to investigate properly (and what I'm saying may well be wrong), but please let me know if you'd like me to look into this further (e.g. make an MVP to demonstrate).
Thanks!
Please see #12 for more details. Tl;dr it seems like we can get a ~6.5x speedup on inference (at least on an L40S with the case of a 32x32 grid with 6 timesteps, one channel and batch size 4) by changing one line of code.
I'm using walrus for a personal project so unfortunately I don't have much time to investigate properly (and what I'm saying may well be wrong), but please let me know if you'd like me to look into this further (e.g. make an MVP to demonstrate).
Thanks!