docs: note the call overhead of a non-LTO precompiled library

Release build, trivial bound functions, against an LTO header-only
module: AppleClang arm64 adds 2-4 ns per call (7-11%), GCC 15 Linux
aarch64 adds 1-2 ns (1-7%). With LTO on the library, both are within a
few percent.

Assisted-by: ClaudeCode:claude-opus-5-5
1 file changed