{"id":13563700,"url":"https://github.com/autonomousvision/stylegan-xl","last_synced_at":"2025-05-16T12:11:42.363Z","repository":{"id":37349919,"uuid":"454535315","full_name":"autonomousvision/stylegan-xl","owner":"autonomousvision","description":"[SIGGRAPH'22] StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets","archived":false,"fork":false,"pushed_at":"2024-06-24T19:12:57.000Z","size":14239,"stargazers_count":978,"open_issues_count":44,"forks_count":115,"subscribers_count":37,"default_branch":"main","last_synced_at":"2025-04-09T07:04:52.730Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/autonomousvision.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-02-01T20:11:25.000Z","updated_at":"2025-03-24T07:42:37.000Z","dependencies_parsed_at":"2024-08-01T13:29:42.031Z","dependency_job_id":null,"html_url":"https://github.com/autonomousvision/stylegan-xl","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/autonomousvision%2Fstylegan-xl","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/autonomousvision%2Fstylegan-xl/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/autonomousvision%2Fstylegan-xl/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/autonomousvision%2Fstylegan-xl/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/autonomousvision","download_url":"https://codeload.github.com/autonomousvision/stylegan-xl/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254527099,"owners_count":22085919,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-01T13:01:22.374Z","updated_at":"2025-05-16T12:11:42.342Z","avatar_url":"https://github.com/autonomousvision.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"\u003cimg src=\"media/banner.png\"\u003e\n\n\n#### [[Project]](https://sites.google.com/view/stylegan-xl/)    [[PDF]](https://arxiv.org/abs/2202.00273)    [![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-blue)](https://huggingface.co/spaces/hysts/StyleGAN-XL)\n\n\nThis repository contains code for our SIGGRAPH'22 paper \"StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets\"\n\nby [Axel Sauer](https://axelsauer.com/), [Katja Schwarz](https://katjaschwarz.github.io/), and [Andreas Geiger](http://www.cvlibs.net/).\n\n\nIf you find our code or paper useful, please cite\n```bibtex\n@InProceedings{Sauer2021ARXIV,\n  author    = {Axel Sauer and Katja Schwarz and Andreas Geiger},\n  title     = {StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets},\n  journal   = {arXiv.org},\n  volume    = {abs/2201.00273},\n  year      = {2022},\n  url       = {https://arxiv.org/abs/2201.00273},\n}\n```\n\n|Rank on Papers With Code|  \u0026nbsp;\n :---  |  :---\n [![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-imagenet-32x32)](https://paperswithcode.com/sota/image-generation-on-imagenet-32x32?p=stylegan-xl-scaling-stylegan-to-large-diverse)|[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-cifar-10)](https://paperswithcode.com/sota/image-generation-on-cifar-10?p=stylegan-xl-scaling-stylegan-to-large-diverse)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-imagenet-64x64)](https://paperswithcode.com/sota/image-generation-on-imagenet-64x64?p=stylegan-xl-scaling-stylegan-to-large-diverse)|[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-ffhq-256-x-256)](https://paperswithcode.com/sota/image-generation-on-ffhq-256-x-256?p=stylegan-xl-scaling-stylegan-to-large-diverse)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-imagenet-128x128)](https://paperswithcode.com/sota/image-generation-on-imagenet-128x128?p=stylegan-xl-scaling-stylegan-to-large-diverse)|[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-ffhq)](https://paperswithcode.com/sota/image-generation-on-ffhq?p=stylegan-xl-scaling-stylegan-to-large-diverse)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-imagenet-256x256)](https://paperswithcode.com/sota/image-generation-on-imagenet-256x256?p=stylegan-xl-scaling-stylegan-to-large-diverse)|[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-pokemon-256x256)](https://paperswithcode.com/sota/image-generation-on-pokemon-256x256?p=stylegan-xl-scaling-stylegan-to-large-diverse)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-imagenet-512x512)](https://paperswithcode.com/sota/image-generation-on-imagenet-512x512?p=stylegan-xl-scaling-stylegan-to-large-diverse)|[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/stylegan-xl-scaling-stylegan-to-large-diverse/image-generation-on-pokemon-1024x1024)](https://paperswithcode.com/sota/image-generation-on-pokemon-1024x1024?p=stylegan-xl-scaling-stylegan-to-large-diverse)\n\n## Related Projects ##\n- Projected GANs Converge Faster (NeurIPS'21) \u0026nbsp;-\u0026nbsp; [Official Repo](https://github.com/autonomousvision/projected_gan) \u0026nbsp;-\u0026nbsp; [![Projected GAN Quickstart](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/gist/xl-sr/757757ff8709ad1721c6d9462efdc347/projected_gan.ipynb)\n- StyleGAN-XL + CLIP (Implemented by CasualGANPapers) \u0026nbsp;-\u0026nbsp;  [![StyleGAN-XL + CLIP](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/CasualGANPapers/unconditional-StyleGANXL-CLIP/blob/main/StyleganXL%2BCLIP.ipynb)\n- StyleGAN-XL + CLIP (Modified by Katherine Crowson to optimize in W+ space) \u0026nbsp;-\u0026nbsp; [![StyleGAN-XL + CLIP](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/drive/1ZEnJE-EUnh-aCXJbu0kVhi8_Qdi2BV-S)\n\n## Requirements ##\n- 64-bit Python 3.8 and PyTorch 1.9.0 (or later). See https://pytorch.org for PyTorch install instructions.\n- CUDA toolkit 11.1 or later.\n- GCC 7 or later compilers. The recommended GCC version depends on your CUDA version; see for example, CUDA 11.4 system requirements.\n- If you run into problems when setting up the custom CUDA kernels, we refer to the [Troubleshooting docs](https://github.com/NVlabs/stylegan3/blob/main/docs/troubleshooting.md#why-is-cuda-toolkit-installation-necessary) of the original StyleGAN3 repo and the following issues: https://github.com/autonomousvision/stylegan_xl/issues/23.\n- Windows user struggling installing the env might find https://github.com/autonomousvision/stylegan_xl/issues/10\n  helpful.\n- Use the following commands with Miniconda3 to create and activate your PG Python environment:\n  - ```conda env create -f environment.yml```\n  - ```conda activate sgxl```\n\n## Data Preparation ##\nFor a quick start, you can download the few-shot datasets provided by the authors of [FastGAN](https://github.com/odegeasslbc/FastGAN-pytorch). You can download them [here](https://drive.google.com/file/d/1aAJCZbXNHyraJ6Mi13dSbe7pTyfPXha0/view). To prepare the dataset at the respective resolution, run\n```\npython dataset_tool.py --source=./data/pokemon --dest=./data/pokemon256.zip \\\n  --resolution=256x256 --transform=center-crop\n```\n\nYou need to follow our progressive growing scheme to get the best results. Therefore, you should prepare separate zips for each training resolution. You can get the datasets we used in our paper at their respective websites ([FFHQ](https://github.com/NVlabs/ffhq-dataset), [ImageNet](https://image-net.org/)).\n\n## Training ##\n\n\u003cimg src=\"media/system.png\"\u003e\n\nFor progressive growing, we train a stem on low resolution, e.g., 16\u003csup\u003e2\u003c/sup\u003e pixels. When the stem is finished, i.e., FID is saturating, you can start training the upper stages; we refer to these as superresolution stages.\n\n#### Training the stem\n\nTraining StyleGAN-XL on Pokemon using 8 GPUs:\n\n```\npython train.py --outdir=./training-runs/pokemon --cfg=stylegan3-t --data=./data/pokemon16.zip \\\n    --gpus=8 --batch=64 --mirror=1 --snap 10 --batch-gpu 8 --kimg 10000 --syn_layers 10\n```\n```--batch``` specifies the overall batch size, ```--batch-gpu``` specifies the batch size per GPU. The training loop will automatically accumulate gradients if you use fewer GPUs until the overall batch size is reached.\n\nSamples and metrics are saved in ```outdir```. If you don't want to track metrics, set ```--metrics=none```. You can inspect fid50k_full.json or run tensorboard in ```training-runs/``` to monitor the training progress.\n\nFor a class-conditional dataset (ImageNet, CIFAR-10), add the flag ```--cond True ```. The dataset needs to contain the class labels; see the [StyleGAN2-ADA repo](https://github.com/NVlabs/stylegan2-ada-pytorch) on how to prepare class-conditional datasets.\n\n#### Training the super-resolution stages\nContinuing with pretrained stem:\n```\npython train.py --outdir=./training-runs/pokemon --cfg=stylegan3-t --data=./data/pokemon32.zip \\\n  --gpus=8 --batch=64 --mirror=1 --snap 10 --batch-gpu 8 --kimg 10000 --syn_layers 10 \\\n  --superres --up_factor 2 --head_layers 7 \\\n  --path_stem training-runs/pokemon/00000-stylegan3-t-pokemon16-gpus8-batch64/best_model.pkl\n```\n\n```--up_factor``` allows to train several stages at once, i.e., with ```--up_factor=4``` and a 16\u003csup\u003e2\u003c/sup\u003e stem you can directly train at resolution  64\u003csup\u003e2\u003c/sup\u003e.\n\nIf you have enough compute, a good tactic is to train several stages in parallel and then restart the superresolution stage training once in a while. The current stage will then reload its previous stem's ```best_model.pkl```. Performance can sometimes drop at first because of domain shift, but the superresolution stage quickly recovers and improves further.\n\n#### Training recommendations for datasets other than ImageNet\nThe default settings are tuned for ImageNet. For smaller datasets (\u003c50k images) or well-curated datasets (FFHQ), you can significantly decrease the model size enabling much faster training. Recommended settings are: ```--cbase 16384 --cmax 256 --syn_layers 7``` and for superresolution stages ```--head_layers 4```.\n\nSuppose you want to train as few stages as possible. We recommend training a 32x32 or 64x64 stem, then directly scaling to the final resolution (as described above, you must adjust ```--up_factor``` accordingly). However, generally, progressive growing yields better results faster as the throughput is much higher at lower resolutions. This can be seen in this figure by [Karras et al., 2017](https://arxiv.org/abs/1710.10196):\n\n\u003cp align=\"center\"\u003e\n  \u003cimg width=\"400\" src=\"https://user-images.githubusercontent.com/29833625/162812365-1a718ec1-13f5-4944-b6c5-f63020c817a6.png\"\u003e\n\u003c/p\u003e\n\n\n## Generating Samples \u0026 Interpolations ##\n\u003cimg src=\"media/teaser.png\"\u003e\n\nTo generate samples and interpolation videos, run\n```\npython gen_images.py --outdir=out --trunc=0.7 --seeds=10-15 --batch-sz 1 \\\n  --network=https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/pokemon256.pkl\n```\nand\n```\npython gen_video.py --output=lerp.mp4 --trunc=0.7 --seeds=0-31 --grid=4x2 \\\n  --network=https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/pokemon256.pkl\n```\nFor class-conditional models, you can pass the class index via ```--class```, a index-to-label dictionary for Imagenet can be found [here](https://github.com/autonomousvision/stylegan_xl/blob/main/media/imagenet_idx2labels.txt). For interpolation between classes, provide, e.g., ```--cls=0-31``` to  ```gen_video.py```. The list of classes has to be the same length as ```--seeds```.\n\nTo generate a conditional sample sheet, run\n```\npython gen_class_samplesheet.py --outdir=sample_sheets --trunc=1.0 \\\n  --network=https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet128.pkl \\\n  --samples-per-class 4 --classes 0-32 --grid-width 32\n```\n\nFor ImageNet models, we enable multi-modal truncation (proposed by [Self-Distilled\nGAN](https://self-distilled-stylegan.github.io/)). We generated 600k find 10k cluster centroids via k-means. For a given samples, multi-modal truncation finds the closest centroids and interpolates towards it. To switch from uni-model to multi-modal truncation, pass\n\n\u003csub\u003e`--centroids-path=https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet_centroids.npy`\u003c/sub\u003e\u003cbr\u003e\n\n|No Truncation| Uni-Modal Truncation | Multi-Modal Truncation\n:---  |  :---:  |  :---:\n\u003cimg src=\"media/no_truncation.png\"\u003e | \u003cimg src=\"media/unimodal_truncation.png\"\u003e| \u003cimg src=\"media/multimodal_truncation.png\"\u003e\n\n## Image Inversion ##\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"media/inversion.gif\" width=\"60%\"\u003e\n\u003c/p\u003e\n\nTo invert a given image via latent optimization, and optionally use our reimplementation of [Pivotal Tuning Inversion](https://arxiv.org/abs/2106.05744), run\n\n```\npython run_inversion.py --outdir=inversion_out \\\n  --target media/jay.png \\\n  --inv-steps 1000 --run-pti --pti-steps 350 \\\n  --network=https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet512.pkl\n```\n\nProvide an image via ```target```, it is automatically resized and center-cropped to match the generator network. You do not need to provide a class for ImageNet models, we infer the class of a given sample via a pretrained classifier.\n\n## Image Editing ##\n\u003cimg src=\"media/editing_banner.png\"\u003e\n\nTo use our reimplementation of [StyleMC](https://arxiv.org/abs/2112.08493), and generate the example above, run\n\n```\npython run_stylemc.py --outdir=stylemc_out \\\n  --text-prompt \"a chimpanzee | laughter | happyness| happy chimpanzee | happy monkey | smile | grin\" \\\n  --seeds 0-256 --class-idx 367 --layers 10-30 --edit-strength 0.75 --init-seed 49 \\\n  --network=https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet128.pkl \\\n  --bigger-network https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet1024.pkl\n```\n\nRecommended workflow:\n\n- Sample images via ```gen_images.py```.\n- Pick a sample and use it as the inital image for ```stylemc.py``` by providing ```--init-seed``` and ```--class-idx```.\n- Find a direction in style space via ```--text-prompt```.\n- Finetune ```--edit-strength```, ```--layers```, and amount of ```--seeds```.\n- Once you found a good setting, provide a larger model via ```--bigger-network```. The script still optimizes the direction for the smaller model, but uses the bigger model for the final output.\n\n## Pretrained Models ##\n\nWe provide the following pretrained models (pass the url as `PATH_TO_NETWORK_PKL`):\n\n|Dataset| Res | FID | PATH\n :---  |  ---:  |  ---:  | :---\nImageNet| 16\u003csup\u003e2\u003c/sup\u003e   |0.73|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet16.pkl`\u003c/sub\u003e\u003cbr\u003e\nImageNet| 32\u003csup\u003e2\u003c/sup\u003e   |1.11|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet32.pkl`\u003c/sub\u003e\u003cbr\u003e\nImageNet| 64\u003csup\u003e2\u003c/sup\u003e   |1.52|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet64.pkl`\u003c/sub\u003e\u003cbr\u003e\nImageNet| 128\u003csup\u003e2\u003c/sup\u003e  |1.77|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet128.pkl`\u003c/sub\u003e\u003cbr\u003e\nImageNet| 256\u003csup\u003e2\u003c/sup\u003e  |2.26|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet256.pkl`\u003c/sub\u003e\u003cbr\u003e\nImageNet| 512\u003csup\u003e2\u003c/sup\u003e  |2.42|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet512.pkl`\u003c/sub\u003e\u003cbr\u003e\nImageNet| 1024\u003csup\u003e2\u003c/sup\u003e |2.51|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/imagenet1024.pkl`\u003c/sub\u003e\u003cbr\u003e\nCIFAR10 | 32\u003csup\u003e2\u003c/sup\u003e   |1.85|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/cifar10.pkl`\u003c/sub\u003e\u003cbr\u003e\nFFHQ    | 256\u003csup\u003e2\u003c/sup\u003e  |2.19|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/ffhq256.pkl`\u003c/sub\u003e\u003cbr\u003e\nFFHQ    | 512\u003csup\u003e2\u003c/sup\u003e  |2.23|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/ffhq512.pkl`\u003c/sub\u003e\u003cbr\u003e\nFFHQ    | 1024\u003csup\u003e2\u003c/sup\u003e |2.02|  \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/ffhq1024.pkl`\u003c/sub\u003e\u003cbr\u003e\nPokemon | 256\u003csup\u003e2\u003c/sup\u003e  |23.97| \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/pokemon256.pkl`\u003c/sub\u003e\u003cbr\u003e\nPokemon | 512\u003csup\u003e2\u003c/sup\u003e  |23.82| \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/pokemon512.pkl`\u003c/sub\u003e\u003cbr\u003e\nPokemon | 1024\u003csup\u003e2\u003c/sup\u003e |25.47| \u003csub\u003e`https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/models/pokemon1024.pkl`\u003c/sub\u003e\u003cbr\u003e\n\n## Quality Metrics ##\nPer default, ```train.py``` tracks FID50k during training. To calculate metrics for a specific network snapshot, run\n\n```\npython calc_metrics.py --metrics=fid50k_full --network=PATH_TO_NETWORK_PKL\n```\n\nTo see the available metrics, run\n```\npython calc_metrics.py --help\n```\n\nWe provide precomputed FID statistics for all pretrained models:\n```\nwget https://s3.eu-central-1.amazonaws.com/avg-projects/stylegan_xl/gan-metrics.zip\nunzip gan-metrics.zip -d dnnlib/\n```\n\n### Further Information\nThis repo builds on the codebase of [StyleGAN3](https://github.com/NVlabs/stylegan3) and our previous project [Projected GANs Converge Faster](https://github.com/autonomousvision/projected_gan).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fautonomousvision%2Fstylegan-xl","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fautonomousvision%2Fstylegan-xl","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fautonomousvision%2Fstylegan-xl/lists"}