{"id":15195454,"url":"https://github.com/dnlcrl/deep-residual-networks-pyfunt","last_synced_at":"2025-10-28T01:30:37.207Z","repository":{"id":79843547,"uuid":"53226204","full_name":"dnlcrl/deep-residual-networks-pyfunt","owner":"dnlcrl","description":"Python implementation of \"Deep Residual Learning for Image Recognition\" (http://arxiv.org/abs/1512.03385 - MSRA, winner team of the 2015 ILSVRC and COCO challenges).","archived":false,"fork":false,"pushed_at":"2017-10-20T20:13:35.000Z","size":4814,"stargazers_count":52,"open_issues_count":0,"forks_count":10,"subscribers_count":6,"default_branch":"master","last_synced_at":"2025-02-01T10:23:04.043Z","etag":null,"topics":["accuracy","affine-layer","cifar","deep-residual-learning","image-classification","image-recognition","ipython-notebook","mnist","numpy","python","residual-networks","strider","training"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/dnlcrl.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2016-03-05T22:22:35.000Z","updated_at":"2023-05-19T09:34:42.000Z","dependencies_parsed_at":null,"dependency_job_id":"ea87776c-9a1f-4eb4-be99-00d3de75624a","html_url":"https://github.com/dnlcrl/deep-residual-networks-pyfunt","commit_stats":null,"previous_names":[],"tags_count":1,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dnlcrl%2Fdeep-residual-networks-pyfunt","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dnlcrl%2Fdeep-residual-networks-pyfunt/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dnlcrl%2Fdeep-residual-networks-pyfunt/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/dnlcrl%2Fdeep-residual-networks-pyfunt/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/dnlcrl","download_url":"https://codeload.github.com/dnlcrl/deep-residual-networks-pyfunt/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":238579176,"owners_count":19495510,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["accuracy","affine-layer","cifar","deep-residual-learning","image-classification","image-recognition","ipython-notebook","mnist","numpy","python","residual-networks","strider","training"],"created_at":"2024-09-27T23:23:50.579Z","updated_at":"2025-10-28T01:30:36.654Z","avatar_url":"https://github.com/dnlcrl.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Deep Residual Learning for Image Recognition\n\nImplementation of [\"Deep Residual Learning for Image Recognition\", Kaiming\nHe, Xiangyu Zhang, Shaoqing Ren, Jian Sun](http://arxiv.org/abs/1512.03385) in [PyFunt](https://github.com/dnlcrl/PyFunt) (a simple Python + Numpy DL framework).\n\nAlso inspired by [this implementation in Lua + Torch](https://github.com/gcr/torch-residual-networks).\n\nThe network operates on minibatches of data that have shape (N, C, H, W)\nconsisting of N images, each with height H and width W and with C input\nchannels. It has, like in the reference paper, (6*n)+2 layers,\ncomposed as below:\n\n\t\t\t                                        (image_dim: 3, 32, 32; F=16)\n\t\t\t                                        (input_dim: N, *image_dim)\n\t\t\t INPUT\n\t\t\t    |\n\t\t\t    v\n\t\t\t+-------------------+\n\t\t\t|conv[F, *image_dim]|                    (out_shape: N, 16, 32, 32)\n\t\t\t+-------------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+-------------------------+\n\t\t\t|n * res_block[F, F, 3, 3]|              (out_shape: N, 16, 32, 32)\n\t\t\t+-------------------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+-------------------------+\n\t\t\t|res_block[2*F, F, 3, 3]  |              (out_shape: N, 32, 16, 16)\n\t\t\t+-------------------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+---------------------------------+\n\t\t\t|(n-1) * res_block[2*F, 2*F, 3, 3]|      (out_shape: N, 32, 16, 16)\n\t\t\t+---------------------------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+-------------------------+\n\t\t\t|res_block[4*F, 2*F, 3, 3]|              (out_shape: N, 64, 8, 8)\n\t\t\t+-------------------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+---------------------------------+\n\t\t\t|(n-1) * res_block[4*F, 4*F, 3, 3]|      (out_shape: N, 64, 8, 8)\n\t\t\t+---------------------------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+-------------+\n\t\t\t|pool[1, 8, 8]|                          (out_shape: N, 64, 1, 1)\n\t\t\t+-------------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t+-------+\n\t\t\t|softmax|                                (out_shape: N, num_classes)\n\t\t\t+-------+\n\t\t\t    |\n\t\t\t    v\n\t\t\t OUTPUT\n\nEvery convolution layer has a pad=1 and stride=1, except for the dimension\nenhancning layers which has a stride of 2 to mantain the computational\ncomplexity.\nOptionally, there is the possibility of setting m affine layers immediatley before the softmax layer by setting the hidden_dims parameter, which should be a list of integers representing the numbe of neurons for each affine layer.\n\nEach residual block is composed as below:\n\n\t          Input\n\t             |\n\t     ,-------+-----.\n\tDownsampling      3x3 convolution+dimensionality reduction\n\t    |               |\n\t    v               v\n\tZero-padding      3x3 convolution\n\t    |               |\n\t    `-----( Add )---'\n\t             |\n\t          Output\n\nAfter every layer, a batch normalization with momentum .1 is applied.\n\n\n## Requirements\n\n- [Python 2.7](https://www.python.org/)\n- [numpy](www.numpy.org/)\n- [pyfunt](https://github.com/dnlcrl/PyFunt)\n- [pydatset](https://github.com/dnlcrl/PyDatSet)\n\n\nAfter you get Python, you can get [pip](https://pypi.python.org/pypi/pip) and install all requirements by running:\n\t\n\tpip install -r requirements.txt\n\n## Usage\n\nIf you want to train the network on the CIFAR-10 dataset, simply run:\n\n\tpython train.py --help\n\t\nOtherwise, you have to get the right train.py for MNIST or SFDDD datasets, they are respectively on the mnist and sfddd git branches:\n\n- train.py for MNIST: https://github.com/dnlcrl/PyResNet/blob/mnist/train.py\n\n- train.py for SFDDD: https://github.com/dnlcrl/PyResNet/blob/sfddd/train.py\n\n## Experiments Results\n\nYou can view all the experiments results in the [./docs directory](https://github.com/dnlcrl/PyResNet/tree/master/docs). Main results are shown below:\n\n###  [CIFAR-10](https://www.cs.toronto.edu/~kriz/cifar.html)\n\nbest error: 9.59 % (accuracy: 0.9041) with a 20 layers residual network (n=3):\n\n[![CIFAR-10 results](https://github.com/dnlcrl/PyResNet/blob/master/docs/imgs/cifar.png)](https://github.com/dnlcrl/PyResNet/blob/master/docs/CIFAR-10%20Experiments.ipynb)\n\n[CIFAR-10 Results - iPython notebook](https://github.com/dnlcrl/PyResNet/blob/master/docs/CIFAR-10%20Experiments.ipynb)\n\n###  [MNIST](http://yann.lecun.com/exdb/mnist/)\n\nbest error: 0.36 % (accuracy: 0.9964) with a 32 layers residual network (n=5):\n\n[![MNIST results](https://github.com/dnlcrl/PyResNet/blob/master/docs/imgs/mnistres.png)](https://github.com/dnlcrl/PyResNet/blob/master/docs/MNIST%20Experiments.ipynb)\n\n[MNIST Results - iPython notebook](https://github.com/dnlcrl/PyResNet/blob/master/docs/MNIST%20Experiments.ipynb)\n\n\n###  [SFDDD](https://www.kaggle.com/c/state-farm-distracted-driver-detection)\n\nbest error: 0.25 % (accuracy: 0.9975 %) on a subset (1000 samples) of the train data (~21k images) with a 44 layers residual network (n=7), resizing the images to 64x48, randomly cropping 32x32 images for training and cropping a 32x32 image from the center of the original images for testing. Unfortunately I got more than 2% error on Kaggle's results (composed of ~80k images).\n\n[![SFDDD results](https://github.com/dnlcrl/deep-residual-networks-pyfunt/blob/master/docs/imgs/SFDDD.png)](https://github.com/dnlcrl/PyResNet/blob/master/docs/SFDDD%20Experiments.ipynb)\n\t\n[SFDDD Results - iPython notebook](https://github.com/dnlcrl/PyResNet/blob/master/docs/SFDDD%20Experiments.ipynb)\n\n\u003c!--## TODOs:\n\n- regenerate plots with english labels\n\n- experimentat elastic distortion as data augmentation function on MNIST\n\n- experiment other data augmentation functions on SFDDD\n\n- implementation of the second version of residual networks, as explained in [\"Identity Mappings in Deep Residual Networks\" by Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun](http://arxiv.org/pdf/1603.05027v1.pdf) --\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdnlcrl%2Fdeep-residual-networks-pyfunt","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdnlcrl%2Fdeep-residual-networks-pyfunt","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdnlcrl%2Fdeep-residual-networks-pyfunt/lists"}