{"id":18887125,"url":"https://github.com/thunlp-mt/directquote","last_synced_at":"2025-10-05T02:56:18.000Z","repository":{"id":45285721,"uuid":"411126199","full_name":"THUNLP-MT/DirectQuote","owner":"THUNLP-MT","description":"A Dataset for Direct Quotation Extraction and Attribution in News Articles.","archived":false,"fork":false,"pushed_at":"2021-09-28T03:40:27.000Z","size":2614,"stargazers_count":14,"open_issues_count":1,"forks_count":3,"subscribers_count":3,"default_branch":"main","last_synced_at":"2025-07-18T06:56:02.155Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/THUNLP-MT.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2021-09-28T03:37:53.000Z","updated_at":"2025-05-05T14:35:40.000Z","dependencies_parsed_at":"2022-08-04T02:30:26.732Z","dependency_job_id":null,"html_url":"https://github.com/THUNLP-MT/DirectQuote","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/THUNLP-MT/DirectQuote","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FDirectQuote","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FDirectQuote/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FDirectQuote/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FDirectQuote/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/THUNLP-MT","download_url":"https://codeload.github.com/THUNLP-MT/DirectQuote/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/THUNLP-MT%2FDirectQuote/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":278403319,"owners_count":25981014,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-05T02:00:06.059Z","response_time":54,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-08T07:34:20.333Z","updated_at":"2025-10-05T02:56:17.979Z","avatar_url":"https://github.com/THUNLP-MT.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# DirectQuote - A Dataset for Direct Quotation Extraction and Attribution in News Articles\n\nDirectQuote is a corpus containing 19,760 paragraphs and 10,353 direct quotations manually annotated from online news media.\n\nA _quotation_ is a general notion that covers different kinds of speech, thought, and writing in text (Semino and Short,2004). It is a prominent linguistic device for expressing opinions, statements, and assessments attributed to the speaker (Cappelen and Lepore, 2012). Among all kinds of quotations, the entire content of the _direct quotation_ (O’Keefe et al.,2013) is in quotation marks, which means that what the speaker said is transcribed verbatim.\n\n## Task Definition\nQuotation extractionis defined as extracting reported speech from a third party in the text, also known as reportedspeech extraction. Quotation attribution refers to determining the speaker of the quotation. When annotating speakers, we ensure that valid speakers should be able to belinked to a person entity in a named entity library. Among them, simple patterns are removed to increase the diversity of the corpus. \n\n## Data \n\n\n\u003ctable\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eRegion\u003c/td\u003e\n      \u003ctd\u003eName\u003c/td\u003e\n      \u003ctd\u003eNumbers\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd rowspan=\"5\"\u003eU.S.\u003c/td\u003e\n      \u003ctd\u003eAssociated Press\u003c/td\u003e\n      \u003ctd\u003e438\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eCable News Network\u003c/td\u003e\n      \u003ctd\u003e627\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eAmerican Broadcasting Company\u003c/td\u003e\n      \u003ctd\u003e240\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eNew York Times\u003c/td\u003e\n      \u003ctd\u003e5,642\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eCBS Broadcasting\u003c/td\u003e\n      \u003ctd\u003e4,890\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd rowspan=\"3\"\u003eUK\u003c/td\u003e\n      \u003ctd\u003eBritish Broadcasting Corporation\u003c/td\u003e\n      \u003ctd\u003e926\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eReuters\u003c/td\u003e\n      \u003ctd\u003e5,836\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eThe Guardian\u003c/td\u003e\n      \u003ctd\u003e4,302\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd rowspan=\"2\"\u003eCanada\u003c/td\u003e\n      \u003ctd\u003eThe Globe and Mail\u003c/td\u003e\n      \u003ctd\u003e1,955\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eThe Star\u003c/td\u003e\n      \u003ctd\u003e13,769\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eNew Zealand\u003c/td\u003e\n      \u003ctd\u003eNZ Herald\u003c/td\u003e\n      \u003ctd\u003e115\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd rowspan=\"2\"\u003eAustralia\u003c/td\u003e\n      \u003ctd\u003eAustralian Broadcasting Corporation\u003c/td\u003e\n      \u003ctd\u003e312\u003c/td\u003e\n   \u003c/tr\u003e\n   \u003ctr\u003e\n      \u003ctd\u003eSydney Morning Herald\u003c/td\u003e\n      \u003ctd\u003e93\u003c/td\u003e\n   \u003c/tr\u003e\n\u003c/table\u003e\n\nWe select representative and multiple news sources across the political spectrum, including 13 well-known online news media from five major English-speaking countries. The corpus adopts the format consistent with CoNLL 2003. We use IOB1 format in the corpus. Raw texts are tokenized by whitespace tokenizer. Every word is classified into the following lables:\n\n* `LeftSpeaker` Quotation, the corresponding speaker is in the preceding text\n* `RightSpeaker` Quotation, the corresponding speaker is in the following text\n* `Unknown` Quotation, no corresponding speaker\n* `Speaker` Speaker\n* `Out` Neither\n\n\n## Statistics\n|              | Numbers         |\n| ------------ | --------------- |\n| News Article | 39,153          |\n| Paragraph    | 19,760          |\n| Quotation    | 10,353          |\n| Time         | 2020.09-2021.03 |\n\n\n## Reference\n_DirectQuote: A Dataset for Direct Quotation Extraction and Attribution in News Articles_, Yuanchi Zhang, Yang Liu","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fthunlp-mt%2Fdirectquote","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fthunlp-mt%2Fdirectquote","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fthunlp-mt%2Fdirectquote/lists"}