Real-SWE: benchmarking AI on private enterprise code
Real-SWE releases a benchmark evaluating frontier AI models on licensed private production codebases from real companies, highlighting challenges of proprietary code, business-critical consequences, and company-specific complexity.