LongHarness Benchmark: 68% Is The Ceiling
LongHarness Bench landed on arXiv September 29, 2026, and the best score in the whole paper is 68%.
That's the macro-average accuracy of the strongest model-runtime pairing across four evaluation suites. And it's the number that should recalibrate how you buy agent tooling this year.
If