Measured results of the Top-K experimental selection vs the mmap baseline.
The Top-K implementation exhibited progressive JIT warming, starting slower but eventually beating the baseline. However, to guarantee consistency for the deployment, the backend was reverted to the mmap full-sort baseline prior to deployment.
The results shown for "Top-K selection" represent experimental warming runs that demonstrated higher ultimate throughput. However, the production deployment currently utilizes the "mmap baseline (deployment version)" (full sort) to guarantee exact BM25 ranking consistency across shards.