An open API service indexing awesome lists of open source software.

https://github.com/dcavar/hoosierellipsiscorpus

The Hoosier Ellipsis Corpus
https://github.com/dcavar/hoosierellipsiscorpus

Last synced: 6 months ago
JSON representation

The Hoosier Ellipsis Corpus

Awesome Lists containing this project

README

          

# The Hoosier Ellipsis Corpus (THEC)

Created by [Damir Cavar], 03/18/2024

Last change: [Damir Cavar], 03/18/2024

The Hoosier Ellipsis Corpus is a dataset created and maintained by a team of researchers at the [NLP-Lab](https://nlp-lab.org/).

The current version 1.0 covers the following languages:

- [Arabic](https://github.com/dcavar/thec_ara)
- [Mandarin Chinese](https://github.com/dcavar/thec_cmn)
- Croatian
- [English](https://github.com/dcavar/thec_eng)
- [German](https://github.com/dcavar/thec_deu)
- [Gujarati](https://github.com/dcavar/thec_guj)
- [Hindi](https://github.com/dcavar/thec_hin)
- [Japanese](https://github.com/dcavar/thec_jpn)
- [Kumaoni](https://github.com/dcavar/thec_kfy)
- [Korean](https://github.com/dcavar/thec_kor)
- [Navajo](https://github.com/dcavar/thec_nav)
- Norwegian
- [Polish](https://github.com/dcavar/thec_pol)
- [Russian](https://github.com/dcavar/thec_rus)
- [Spanish](https://github.com/dcavar/thec_spa)
- Swedish
- [Telugu](https://github.com/dcavar/thec_tel)​
- [Ukrainian](https://github.com/dcavar/thec_ukr)​

The following languages are in preparation:

- Bengali
- Bosnian
- Bulgarian
- Hebrew
- Kanada
- Serbian
- Slovak
- Slovenian
- Tamil

Each language dataset is stored in its own repo. Follow the links to download the dataset.

If you have data to contribute to some language, contact us at the [NLP-Lab](https://nlp-lab.org/) or contact [Damir Cavar].

## Citation

Please use the following snippets to cite our work.

```bibtex
@inproceedings{cavar-etal-2024-typology,
title = "The Typology of Ellipsis: A Corpus for Linguistic Analysis and Machine Learning Applications",
author = "Cavar, Damir and Mompelat, Ludovic and Abdo, Muhammad",
editor = "Hahn, Michael and Sorokin, Alexey and Kumar, Ritesh and Shcherbakov, Andreas and Otmakhova, Yulia and Yang, Jinrui and Serikov, Oleg and Rani, Priya and Ponti, Edoardo M. and Murado{\u{g}}lu, Saliha and Gao, Rena and Cotterell, Ryan and Vylomova, Ekaterina",
booktitle = "Proceedings of the 6th Workshop on Research in Computational Linguistic Typology and Multilingual NLP",
month = mar,
year = "2024",
address = "St. Julian's, Malta",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.sigtyp-1.6",
pages = "46--54"
}

@inproceedings{cavar-atal-2004-computing,
author = "Cavar, Damir and Zoran Tiganj and Ludovic Mompelat and Billy Dickson",
title={Computing Ellipsis Constructions: Comparing Classical {NLP} and {LLM} Approaches},
booktitle={2024 Meeting of the Society for Computation in Linguistics (SCiL)},
year={2024}
}
```

[Damir Cavar]: http://damir.cavar.me/ "Damir Cavar"
[Hoosier Ellipsis Corpus]: https://nlp-lab.org/ellipsis/ "Hoosier Ellipsis Corpus"
[the Hoosier Ellipsis Corpus]: https://nlp-lab.org/ellipsis/ "the Hoosier Ellipsis Corpus"
[NLP-Lab]: https://nlp-lab.org/ "NLP-Lab"