|
Download datasets/polycoder.md from codeparrot/code-generation-models: direct link, hf CLI and curl.
- Browser
- Download file 486 Bytes
-
https://huggingface.co/spaces/codeparrot/code-generation-models/resolve/refs%2Fpr%2F13/datasets/polycoder.md
- Command line
-
hf download hf://spaces/codeparrot/code-generation-models@refs/pr/13/datasets/polycoder.md
-
curl -L -o polycoder.md https://huggingface.co/spaces/codeparrot/code-generation-models/resolve/refs%2Fpr%2F13/datasets/polycoder.md
486 Bytes
| The [PolyCoder paper](https://arxiv.org/pdf/2202.13169v3.pdf) gives a nice comparison of existing code models. The authors also trained a code generation model on **249GB** of data, after preprocessing, consisting of popular repositories for 12 popular programming languages with at least 50 stars from GitHub in October 2021. The data used the following preprocessing: | |
| - Exact match deduplication | |
| - Filtering: | |
| - Average line length < 100 tokens | |
| - Maximum line length < 1000 MB |