{"id":17834987,"url":"https://github.com/motykatomasz/pythia-ai-code-completion","last_synced_at":"2025-10-05T16:03:14.881Z","repository":{"id":67540341,"uuid":"413768274","full_name":"motykatomasz/Pythia-AI-code-completion","owner":"motykatomasz","description":"Project reproducing paper: \"Pythia, AI-assisted code completion system\". The project was done for the course \"Machine Learning for Software Engineering\" at TU Delft.","archived":false,"fork":false,"pushed_at":"2021-10-05T10:23:15.000Z","size":134,"stargazers_count":5,"open_issues_count":0,"forks_count":1,"subscribers_count":2,"default_branch":"master","last_synced_at":"2024-10-27T23:14:07.776Z","etag":null,"topics":["abstract-syntax-tree","code-completion","lstm","pytorch"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/motykatomasz.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2021-10-05T10:20:21.000Z","updated_at":"2024-09-30T10:08:30.000Z","dependencies_parsed_at":"2024-03-15T16:16:46.949Z","dependency_job_id":null,"html_url":"https://github.com/motykatomasz/Pythia-AI-code-completion","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/motykatomasz%2FPythia-AI-code-completion","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/motykatomasz%2FPythia-AI-code-completion/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/motykatomasz%2FPythia-AI-code-completion/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/motykatomasz%2FPythia-AI-code-completion/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/motykatomasz","download_url":"https://codeload.github.com/motykatomasz/Pythia-AI-code-completion/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":229786688,"owners_count":18123997,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["abstract-syntax-tree","code-completion","lstm","pytorch"],"created_at":"2024-10-27T20:16:04.395Z","updated_at":"2025-10-05T16:03:09.861Z","avatar_url":"https://github.com/motykatomasz.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Group 3: Code Completion\n**Paper replicating:** [Pythia: AI-assisted Code Completion System](https://arxiv.org/abs/1912.00742)  \n**Dataset:** [150k Python Dataset](https://www.sri.inf.ethz.ch/py150)\n\n#### Dataset statistics\nPython150 marks 100k files as training files and 50k files as evaluation/test files. Using a [deduplication tool](https://github.com/saltudelft/CD4Py) the original files were deduplicated resulting in **84728** training files and **42372** test files.  \n\nIn the preprocessing phase files which couldn't be parsed to an AST (e.g. because the Python version was too old) were removed. This reduced the training files to **75183** AST's. From these AST's a vocabulary of tokens is built using a threshold of `20`. This means that tokens which occur in the training set more than 20 times will be added to the vocabulary. This resulted in a vocabulary of size **43853**. \n\n## Experiments\nShared parameters:\n```\nbatch size: 64\nembed dimension: 150\nhidden dimension: 500\nnum LSTM layers: 2\nlookback tokens: 100\nnorm clipping: 10\ninitial learning rate: 2e-3\nlearning rate schedule: decay of 0.97 every epoch\nepochs: 15\n```\n\nDeployed code: [experiments release](https://github.com/serg-ml4se-2020/group3-code-completion/releases/tag/experiments)\n\n**Data used:**  \nTraining set (1970000 items): [download here](https://drive.google.com/file/d/1cARlxinp1y7bQqXBWbVLgJUdQ9lyJi9g/view?usp=sharing)  \nValidation set (227255 items): [download here](https://drive.google.com/file/d/1EObNu3m24id-t60nK8hxqSZnhhsmWQRu/view?usp=sharing)  \nEvaluation set (911213 items): [download here](https://drive.google.com/file/d/1wh1viWN7q3XNJGljr6tWTubMW6bYTRaj/view?usp=sharing)  \nVocabulary (size: 43853): [download here](https://drive.google.com/file/d/132HLLacrL_lWZfnYmRORtldkUpMb1Jgh/view?usp=sharing)\n\n### Experiment 1 | Regularization - L2 regularizer\nL2 parameter of `1e-6` [(also done here)](https://arxiv.org/abs/1701.06548).\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        46.61%  |        71.67%  |\n| Evaluation set |        47.89%  |        69.76%  |\n\n![plot](results/experiment1/experiment1.svg)\n\nResulting model: [final_model_experiment_1](https://drive.google.com/file/d/1zZdN5fg3bHD1e_hp1WMOe4J_EM_vAfz3/view?usp=sharing)\n### Experiment 2 | Regularization - Dropout\n\nDropout parameter of `0.8` (based on Pythia).\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        38.53%       |        63.31%       |\n| Evaluation set |         39.37%      |        61.15%       |\n\n![plot](results/experiment2/experiment2.svg)\n\nResulting model: [final_model_experiment_2](https://drive.google.com/file/d/1rOGaq_FIOBhzKCppi8Tr8bZdhH65Nt1K/view?usp=sharing)\n### Experiment 3 | No regularization\nNo L2, dropout or weighted loss.\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        43.03%       |        67.37%       |\n| Evaluation set |         46.63%      |        68.67%       |\n\n![plot](results/experiment3/experiment3.svg)\n\nResulting model: [final_model_experiment_3](https://drive.google.com/file/d/1x6YWdyZLNowmyWGMcpQdvHvfWzGzxDyi/view?usp=sharing)\n### Experiment 4 | Regularization - Weighted loss + L2\n\nIncludes a weighted loss + L2  (`1e-6`).\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        40.53%       |        63.62%       |\n| Evaluation set |        41.31%       |        60.84%       |\n\n![plot](results/experiment4/experiment4.svg)\n\nResulting model: [final_model_experiment_4](https://drive.google.com/file/d/1icevmNV0Bj1sdI2UMUG25L8e6iCp-qmd/view?usp=sharing)\n### Experiment 5 | bi-directional LSTM\n\nUsing a bi-directional LSTM instead of uni-directional. Also includes L2 regularizer  (`1e-6`). \n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        48.60%       |        71.49%       |\n| Evaluation set |        49.87%       |        70.11%       |\n\n\n![plot](results/experiment5/experiment5.svg)\n\nResulting model: [final_model_experiment_5](https://drive.google.com/file/d/1hM3HE-xNJLdCCc0gIO7CmC-Bglg9Or61/view?usp=sharing)\n\n### Experiment 6 | attention\n\nUsing an attention mechanism. Also includes L2 regularizer. \n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        51.10%       |        73.95%       |\n| Evaluation set |        52.90%       |        72.86%       |\n\n\n![plot](results/experiment6/experiment6.svg)\n\nResulting model: [final_model_experiment_6](https://drive.google.com/file/d/14S0b3CJlU3lF3Z9mByCAEMSFKx2xygd6/view?usp=sharing)\n\n### Experiment 7 | attention\n\nUsing an attention mechanism. Also includes L2 regularizer (`1e-6`) and a (lower) dropout (`0.4`). \n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        51.51%       |        74.17%       |\n| Evaluation set |        53.70%       |        73.22%       |\n\n![plot](results/experiment7/experiment7.svg)\n\nResulting model: [final_model_experiment_7](https://drive.google.com/file/d/1aYZk2WrjSCKca-V5h_dWYVmU4nABN2X6/view?usp=sharing)\n### Experiment 8 | Regularizer - (low) dropout\n\nIncludes L2 regularizer (`1e-6`) and a (lower) dropout (`0.4`)\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        47.31%       |        71.89%       |\n| Evaluation set |        48.56%       |        70.71%       |\n\n![plot](results/experiment8/experiment8.svg)\nResulting model: [final_model_experiment_8](https://drive.google.com/file/d/12OL-xGfMvmEfOdB78KyjQhDkKjjhyLgS/view?usp=sharing)\n\n### Experiment 9 | Attention different\n\nIncludes L2 regularizer (`1e-6`) and a (lower) dropout (`0.4`)\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        50.67%       |        73.38%       |\n| Evaluation set |        52.69%%       |        72.36%       |\n\n\n![plot](results/experiment9/experiment9.svg)\nResulting model: [final_model_experiment_9](https://drive.google.com/file/d/1kJiJ2ELB6DdyYmTX4VrTGoU2SBdE5JSE/view?usp=sharing)\n\n### Experiment 10 | Attention final\n\nIncludes L2 regularizer (`1e-6`) and a (lower) dropout (`0.4`). Runs for 30 epochs.\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        52.20%       |        75.12%       |\n| Evaluation set |        54.80%       |        74.51%       |\n\n\n![plot](results/experiment10/experiment10.svg)\nResulting model: [final_model_experiment_10](https://drive.google.com/file/d/1bGv8a_2Qhh1urN8r0S0wUluL8KYvDt4I/view?usp=sharing)\n\n### Experiment 11 | Best regularizer\n\nIncludes L2 regularizer (`1e-6`) and a (lower) dropout (`0.4`).\n\n|                | **Top-1 accuracy** | **Top-5 accuracy** |\n|----------------|----------------|----------------|\n| Validation set |        47.26%       |        71.96%       |\n| Evaluation set |        48.51%       |        70.60%       |\n\n\n![plot](results/experiment11/experiment11.svg)\nResulting model: [final_model_experiment_11](https://drive.google.com/file/d/1AatStOh3P-a3e7qWCYyK0I_FiR-AYQrz/view?usp=sharing)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmotykatomasz%2Fpythia-ai-code-completion","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmotykatomasz%2Fpythia-ai-code-completion","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmotykatomasz%2Fpythia-ai-code-completion/lists"}