The 50% is average overhead and basically entirely due to context switch overhead and execution serialization. If the execution is entirely in userspace, rr's overhead is basically 0. I don't think PT will help here. That said, rr's performance on Intel chips is entirely acceptable for single threaded code. The big asks would be other architectures, or as roc mentioned, something like QuickRec for efficient multi-threaded recording.