SWE-bench

1 report
SWE-bench is an artificial intelligence benchmark that compares systems on defined tasks, datasets, metrics, and testing conditions. Evaluation relies on benchmark dataset, test split, and data contamination, including the costs, limitations, and tradeoffs hidden by a single headline metric.

The evidence base around SWE-bench is examined through baseline systems, together with benchmark dataset and test split. Complementary views of test split come from held-out evaluations and independent reproduction, but the conclusion remains bounded because the comparison must account for the fact that a single score may not measure robustness, safety, or practical usefulness.