{"id":13711303,"url":"https://github.com/benkrause/dynamic-evaluation","last_synced_at":"2025-05-06T20:32:43.290Z","repository":{"id":201731325,"uuid":"105284800","full_name":"benkrause/dynamic-evaluation","owner":"benkrause","description":"Dynamic evaluation for pytorch language models, now includes hyperparameter tuning","archived":false,"fork":false,"pushed_at":"2017-11-24T15:27:52.000Z","size":9,"stargazers_count":105,"open_issues_count":1,"forks_count":21,"subscribers_count":4,"default_branch":"master","last_synced_at":"2024-08-03T23:23:25.334Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"bsd-2-clause","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/benkrause.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2017-09-29T14:57:31.000Z","updated_at":"2024-04-18T05:34:05.000Z","dependencies_parsed_at":null,"dependency_job_id":"2dacde81-0087-4953-8f8c-5c5fcca781be","html_url":"https://github.com/benkrause/dynamic-evaluation","commit_stats":null,"previous_names":["benkrause/dynamic-evaluation"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benkrause%2Fdynamic-evaluation","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benkrause%2Fdynamic-evaluation/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benkrause%2Fdynamic-evaluation/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/benkrause%2Fdynamic-evaluation/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/benkrause","download_url":"https://codeload.github.com/benkrause/dynamic-evaluation/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":224528396,"owners_count":17326358,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-02T23:01:06.758Z","updated_at":"2024-11-13T21:31:45.544Z","avatar_url":"https://github.com/benkrause.png","language":"Python","funding_links":[],"categories":["Machine Learning and Deep Learning"],"sub_categories":[],"readme":"#### Dynamic evaluation for pytorch language models as implemented in [Dynamic Evaluation of Neural Sequence Models](https://arxiv.org/abs/1709.07432). \n\n#### Requirements: python 3 (tested in 3.5, 3.6), pytorch (tested in 0.1.12, 0.2)\n\n#### Instructions for use:  \n\n1. Train a language model using an existing repository, such as the [pytorch language modeling tutorial](https://github.com/pytorch/examples/tree/master/word_language_model) . This should save a .pt file with the trained model\n\n2. Copy the file dynamiceval.py into the repository\n\n3. Run dynamic evaluation with: `python dynamiceval.py --model modelname.pt`\n\n#### AWD-LSTM\n\nTo replicate results in paper for AWD-LSTM + dynamic eval, train the language model using the [Salesforce AWD-LSTM repository](https://github.com/salesforce/awd-lstm-lm). We used the original codebase from this repository, with the goal of exact replication of results from their paper (which we failed to achieve). The default settings and hyper-parameters for dynamiceval.py are tuned for for AWD-LSTM + dynamic eval on PTB. \n\n#### AWD-QRNN\n\nThis code also supports the [pytorch QRNN](https://github.com/salesforce/pytorch-qrnn) with the --QRNN option. AWD-QRNN + dynamic eval obtains very similar results to AWD-LSTM + dynamic eval, and is much faster to train and evaluate. Training an AWD-QRNN on PTB using the Salesforce AWD-LSTM repository, and running dynamic eval with the default settings gives a test perplexity of 50.5. Increasing the sequence segment length from 5 to 20 runs 3x faster (1 minute vs. 3 minutes on PTB), and gives validation (use --val flag) and test perplexities of `51.4/50.5` with the following arguments:\n\n`python dynamiceval.py --model PTB.pt --QRNN --lr 0.00012 --lamb 0.02 --bptt 20`\n\nAWD-QRNN trained on wikitext-2 gives validation (use --val flag) and test perplexities of `45.9/44.0` with the following arguments:\n\n`python dynamiceval.py --model WT2.pt --QRNN --lr 0.00012 --lamb 0.008 --bptt 20 --data data/wikitext-2`\n\n#### Hyper-parameter search\n\nTo get stronger results with any other model or dataset, you can run with:\n\n`python dynamiceval.py --model modelname.pt --grid`\n\nThis will do a hyper-parameter search on the validation set, takes a few hours on PTB with LSTM and default settings, can be much faster with QRNN and/or larger --bptt. If the default model size/settings are changed too much, the hyper-parameters in `lrlist` and `lamblist` may need to be changed for best results. If you want to do a faster search, you can try running with --gridfast to use a subset of the validation set, or you can reduce the number of elements in lamblist (tuning lr is more important).\n\n#### Command line arguments:\n\n\n`--model` (required)    -filename of the trained model to be evaluated\n\n`--data`    -location of the data corpus\n\n`--grid`    -hyper-parameter grid search over lambda and eta, gives both valid and test error\n\n`--gridfast`    -same as grid, but only uses first 30k validation tokens for search\n\n`--val `   -measure validation error instead of test error  \n\n`--gpu`    -specify a gpu device, uses device 0 by default (set negative for cpu)\n\n`--QRNN`    -apply dynamic eval to a QRNN\n\n`--bptt`    -sequence segment length for dynamic eval, also used for gradient statistics on training data\n\n`--batch_size`    -batch size for gradient statistics on training data\n\n`--lr`    -learning rate eta (ignored if --grid is set)\n\n`--lamb`    -decay rate lambda (ignored if --grid is set)\n\n`--epsilon`    -stabilization parameter epsilon\n\n`--max_batches`  -max number of batches for training gradient statistics (-1 uses full training set)\n\n`--oldhyper`\n\n-The original version code inadvertently scaled a couple of terms differently from the equations described in the paper, which affects the hyper-parameters. The code was changed to reflect the paper equations, and this flag applies a hyper-parameter transformation in a way that accounts for this change. If you run this version of the code with the --oldhyper flag, it is equivalent to running the old version of the code. This will also print out hyper-parameter values that can be used with the new version of the code (without this flag) to obtain the same results. Some previous results that report hyper-parameters with the old code would require applying this hyper-parameter scaling to achieve the same results with this code. The following replicates the exact settings used to obtain results for AWD-LSTM in the paper for PTB:\n\n`python dynamiceval.py --model PTB.pt --lr 0.002 --lamb 0.02 --epsilon 0.001 --oldhyper`\n\nand for wikitext-2:\n\n`python dynamiceval.py --model WT2.pt --lr 0.002 --lamb 0.02 --epsilon 0.001 --oldhyper --data data/wikitext-2`\n\n\nNote that while the original hyper-parameters are the same for Wikitext-2 and PTB, the new/scaled hyper-parameters are a little different.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbenkrause%2Fdynamic-evaluation","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fbenkrause%2Fdynamic-evaluation","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fbenkrause%2Fdynamic-evaluation/lists"}