Repository-Level Dynamic Benchmarking Ends the Memorization Game
Code2Bench-2505 spun 880 recent Python projects into 1,163 benchmark tasks, and repository-level dynamic benchmarking stopped being a slide-deck phrase. The tasks came out of an automated pipeline, the test suites were synthesized with property-based testing. And the passing bar was execution rather than anyone's opinion of correct