{"id":50632279,"url":"https://github.com/scaleapi/SWE-bench_Pro-os","last_synced_at":"2026-06-23T19:00:47.162Z","repository":{"id":315647124,"uuid":"1051347649","full_name":"scaleapi/SWE-bench_Pro-os","owner":"scaleapi","description":"SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?","archived":false,"fork":false,"pushed_at":"2026-05-18T18:59:22.000Z","size":34127,"stargazers_count":389,"open_issues_count":25,"forks_count":63,"subscribers_count":5,"default_branch":"main","last_synced_at":"2026-05-18T20:40:38.309Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/scaleapi.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-09-05T20:42:53.000Z","updated_at":"2026-05-18T18:59:27.000Z","dependencies_parsed_at":"2025-10-21T22:27:54.803Z","dependency_job_id":"813f0619-ae71-41a9-9d1f-e19222f88656","html_url":"https://github.com/scaleapi/SWE-bench_Pro-os","commit_stats":null,"previous_names":["scaleapi/swe-bench_pro-os"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/scaleapi/SWE-bench_Pro-os","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scaleapi%2FSWE-bench_Pro-os","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scaleapi%2FSWE-bench_Pro-os/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scaleapi%2FSWE-bench_Pro-os/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scaleapi%2FSWE-bench_Pro-os/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/scaleapi","download_url":"https://codeload.github.com/scaleapi/SWE-bench_Pro-os/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/scaleapi%2FSWE-bench_Pro-os/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34702919,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-23T02:00:07.161Z","response_time":65,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2026-06-06T23:00:22.904Z","updated_at":"2026-06-23T19:00:47.150Z","avatar_url":"https://github.com/scaleapi.png","language":"Python","funding_links":[],"categories":["Catalog","Coding \u0026 Repository Environments"],"sub_categories":["Evaluation Harnesses \u0026 Benchmarks"],"readme":"## SWE-Bench Pro\n\nCode and data for the following works:\n* \u003ca href=\"https://static.scale.com/uploads/654197dc94d34f66c0f5184e/SWEAP_Eval_Scale%20(9).pdf\"\u003eSWE-bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?\u003c/a\u003e\n\n* HuggingFace: \u003ca href=\"https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro\"\u003ehttps://huggingface.co/datasets/ScaleAI/SWE-bench_Pro\u003c/a\u003e\n\n* Public Leaderboard: \u003ca href=\"https://scale.com/leaderboard/swe_bench_pro_public\"\u003ehttps://scale.com/leaderboard/swe_bench_pro_public\u003c/a\u003e\n\n* Commercial (Private) Leaderboard: \u003ca href=\"https://labs.scale.com/leaderboard/swe_bench_pro_private\"\u003ehttps://labs.scale.com/leaderboard/swe_bench_pro_private\u003c/a\u003e\n\n## News\n\n(05/18) We have identified some issues with the leaderboard and are currently working on addressing them. \n\n(2/9) We have removed some unit tests which were outdated (e.g. required the year 2025) or were previously not intended to be included. \n\n(1/7) We have fixed an issue with tutao instances where they take a long time to eval. The relevant run scripts are updated.\n\n(10/28) We added mini-swe-agent! Results are comparable to SWE-Agent for Sonnet 4.5. Feel free to give it a shot. (credit @miguelrc-scale)\n\n(10/28) We have the SWE-Agent scaffold to reproduce results and a step-by-step guide below. We have confirmed that this reproduces the Sonnet 4.5 results. (credit @18vijayb)\n\n(10/3) We have updated results without cap limit here: https://scaleapi.github.io/SWE-bench_Pro-os/\n\n## Overview\nSWE-Bench Pro is a challenging benchmark evaluating LLMs/Agents on long-horizon software engineering tasks.\nGiven a *codebase* and an *issue*, a language model is tasked with generating a *patch* that resolves the described problem.\n\nThe dataset is inspired from SWE-Bench: https://github.com/SWE-bench/SWE-bench\n\nTo access SWE-bench Pro, copy and run the following code:\n```python\nfrom datasets import load_dataset\nswebench = load_dataset('ScaleAI/SWE-bench_Pro', split='test')\n```\n\n## Installation\n\n### 1. Install Python Dependencies\n\n```bash\npip install -r requirements.txt\n```\n\n### 2. Install Docker\n\nSWE-bench Pro uses Docker for reproducible evaluations.\n\nFollow the instructions in the [Docker setup guide](https://docs.docker.com/engine/install/) to install Docker on your machine.\nIf you're setting up on Linux, we recommend seeing the [post-installation steps](https://docs.docker.com/engine/install/linux-postinstall/) as well.\n\n### 3. Configure Modal (Recommended) (or use local docker [Beta])\n\n```bash\nmodal setup  # Follow the prompts to generate your token\n```\n\nAfter running, verify your credentials in `~/.modal.toml`:\n```\ntoken_id = \u003ctoken id\u003e\ntoken_secret = \u003ctoken secret\u003e\nactive = true\n```\n\nBeta: Local Docker. No additional setup needed. Use the `--use_local_docker` flag when running evaluations.\n\n## Docker Images\n\nWe provide prebuilt Docker images for each instance on Docker Hub:\n\n**Repository:** https://hub.docker.com/r/jefzda/sweap-images\n\n### Finding the Correct Image\n\nEach instance in the HuggingFace dataset has a `dockerhub_tag` column containing the Docker tag for that instance. You can access it directly:\n\n```python\nfrom datasets import load_dataset\n\ndataset = load_dataset('ScaleAI/SWE-bench_Pro', split='test')\n\n# Get the Docker image for a specific instance\nfor row in dataset:\n    instance_id = row['instance_id']\n    docker_tag = row['dockerhub_tag']\n    full_image = f\"jefzda/sweap-images:{docker_tag}\"\n    print(f\"{instance_id} -\u003e {full_image}\")\n```\n\n**Important:** Bash runs by default in our images. When running these images, you should not manually invoke bash. See https://github.com/scaleapi/SWE-bench_Pro-os/issues/6\n\n## Usage\n\n### 1. Generate Patches\nGenerate patch predictions using your harness of choice. \n\nFor generating patches using SWE-agent, see the [SWE-agent git submodule](./SWE-agent/) (note: you will have to use this as a git submodule. See [official git documentation](https://git-scm.com/book/en/v2/Git-Tools-Submodules) for details). The submodule contains detailed instructions to\n- Set up SWE-agent for patch generation\n- Run SWE-agent on SWE-Bench Pro instances\n- Configure model parameters and turn limits\n\nThe output will be `.pred` files containing model-generated patches for each instance.\n\n### 2. Gather Patches\nAfter generating patches, use the `gather_patches.py` helper script to collect all patches into a single JSON file for evaluation:\n\n```bash\npython helper_code/gather_patches.py \\\n    --directory \u003cpath_to_pred_files\u003e \\\n    --prefix \u003cmodel_name\u003e \\\n    --output \u003coutput_file\u003e.json\n```\n\n**Parameters:**\n- `--directory`: Directory containing instance folders with `.pred` files (e.g., from SWE-agent output or downloaded trajectories)\n- `--prefix`: Prefix identifier for your model/run (e.g., \"gpt4\", \"claude-sonnet\", \"sample1\")\n- `--output`: Output JSON file path\n\n**Example:**\n```bash\npython helper_code/gather_patches.py \\\n    --directory swe_bench_pro_results/sample1 \\\n    --prefix sample1 \\\n    --output sample1_patches.json\n```\n\nThis will create a JSON file in the format expected by the evaluation script:\n```json\n[\n  {\n    \"instance_id\": \"instance_...\",\n    \"patch\": \"diff --git ...\",\n    \"prefix\": \"sample1\"\n  }\n]\n```\n\n### 3. Evaluate Patches\n\nEvaluate patch predictions on SWE-Bench Pro:\n\n```bash\npython swe_bench_pro_eval.py \\\n    --raw_sample_path=swe_bench_pro_full.csv \\\n    --patch_path=\u003cyour_patches\u003e.json \\\n    --output_dir=\u003coutput_directory\u003e \\\n    --scripts_dir=run_scripts \\\n    --num_workers=100 \\\n    --dockerhub_username=jefzda\n```\n\nYou can test with the gold patches, which are in the HuggingFace dataset. There is a helper script in `helper_code` which can extract the gold patches into the required JSON format.\n\n## Reproducing Leaderboard Results\n\nTo reproduce leaderboard results end-to-end, follow the following steps:\n\n1. Complete setup in the `SWE-agent` submodule. We recommend to use the Docker image to run the scaffold, via `just`.\n2. Run the scaffold. We have included an example for Claude Sonnet 4.5 (claude.yaml) but feel free to use any model. It also supports `vllm` for local models. Note that we recommend using the DockerHub images rather than building the Docker images from scratch. You can also execute it locally without Modal.\n3. Compile predictions with helper_code/gather_patches.py.\n4. Run the evaluation script `swe_bench_pro_eval.py` to run the evaluation script.\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fscaleapi%2FSWE-bench_Pro-os","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fscaleapi%2FSWE-bench_Pro-os","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fscaleapi%2FSWE-bench_Pro-os/lists"}