{"id":21478982,"url":"https://github.com/mjpost/acl2014-handbook","last_synced_at":"2026-01-03T14:37:44.817Z","repository":{"id":17807906,"uuid":"20696952","full_name":"mjpost/acl2014-handbook","owner":"mjpost","description":"Code and data used to produce the ACL 2014 handbook","archived":false,"fork":false,"pushed_at":"2014-10-13T02:19:00.000Z","size":11000,"stargazers_count":2,"open_issues_count":0,"forks_count":1,"subscribers_count":2,"default_branch":"master","last_synced_at":"2025-02-06T02:27:19.160Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"TeX","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/mjpost.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2014-06-10T18:34:55.000Z","updated_at":"2020-04-17T17:28:50.000Z","dependencies_parsed_at":"2022-09-02T13:01:31.902Z","dependency_job_id":null,"html_url":"https://github.com/mjpost/acl2014-handbook","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mjpost%2Facl2014-handbook","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mjpost%2Facl2014-handbook/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mjpost%2Facl2014-handbook/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mjpost%2Facl2014-handbook/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/mjpost","download_url":"https://codeload.github.com/mjpost/acl2014-handbook/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243998707,"owners_count":20381207,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-23T11:20:23.678Z","updated_at":"2026-01-03T14:37:44.771Z","avatar_url":"https://github.com/mjpost.png","language":"TeX","funding_links":[],"categories":[],"sub_categories":[],"readme":"If you're reading this, it's likely because you've been asked to\nassemble an *ACL handbook. Congratulations!  \n\nYou should start your preparation by clearing the entire week before\nthe printed deadline --- really.  The process of assembling the\nhandbook is a little bit of scripting and auto-generation followed by\nan immense manual effort. You will take the proceedings from the\nvarious workshops and pieces of the main conference, use the scripts\nin the `scripts/` subdirectory to generate some LaTeX, and then\nhand-assemble and munge everything until it is complete. Then you will\npore through it finding and correcting errors. You will find many of\nthese errors after you have uploaded the document to the printer, and\nthe last errors will not be found until after printing has begun.\n\nOn the positive side, it can be a very satisfying endeavor, and you\nhave the consolation that you are producing something very useful to\nall the conference attendees, and also something you will likely be\ncredited for. Unlike the bulk of the work of conference organization,\nyour labor is visible and immediate.\n\nThis guide is intended for Publications Chairs (who I think should be\ndoing the job of assembling the handbook) and --- assuming my advice\nis not heeded --- the Handbook chair. I hope that the information\nprovided here will be of some aid to you.\n  \n# Contents of a good handbook\n\nA good conference handbook should have the following\n\n- A nice, customized cover\n- Letters from the General Chair and PC Chairs\n- Schedules for the main conference, workshops, and any co-located conferences\n- Overviews for each day that summarize the main events\n- An \"at-a-glance\" overview page for each parallel session listing times, \n  locations, titles, and authors\n- Abstracts for all papers presented in the main conference\n- Full pages for special events (the banquet, business meeting, special\n  awards and ceremonies)\n- A perfect index, pointing to any and every mention of a name within\n  the document, and containing no duplicate names (with spelling variations)\n- A local guide with restaurants, area attractions, etc\n- Sponsors and ads\n\n# Some terminology\n\n- The conference **proceedings** is the book containing all the\n  conference's accepted papers and hosted in the\n  [ACL anthology](http://aclweb.org/anthology). This is also sometimes\n  called the **book**. Typically, separate books are produced for\n  short and long papers, the student research workshop, and\n  demonstrations.\n  \n- **ACLPUB** is the tool used to assemble the proceedings for a\n  conference. This was, I believe, written in the early 2000s by Jason\n  Eisner and Philipp Koehn. It usually consists of a directory named\n  `proceedings/`, within which are papers, paper metadata, files used to\n  generate the proceedings PDF, and, most importantly for handbook,\n  the `order` file, which specifies the conference schedule.\n  \n- **START** is the online system used to manage paper submission,\n  reviewing, acceptance / rejection, and so on. It knows about ACLPUB, and provides interfaces\n  for editing the `order` file (directly or via their ScheduleMaker\n  tool) and for generating the proceedings.tgz files.\n  \n- **Softconf** is the company that develops START\n\n- **Subconferences**. Conferences hosted in START are assigned a top-level\n  directory, e.g., `softconf.com/acl2014`. Beneath this are individual\n  \"subconferences\" representing the workshops and also pieces of the\n  the main conference. Main conference papers are typically hosted at\n  `papers/`, short papers at `shortpapers`, and so on for \"demos\",\n  \"srw\", \"tutorials\", and \"tacl\". Different sessions are\n  self-contained and isolated and do not know about each other or have\n  the ability to reference each other. The main conference spreads\n  across multiple sessions, while workshops are typically contained in\n  a single one.\n\n# Main conference versus workshops\n\nThere are two main types of events you have to format for: the main\nconference, and workshops. The technical distinction pertinent to\nproducing the handbook is that \n\n- Main conferences are multi-track and require you to assemble\n  schedules from multiple sessions whose schedules may be interleaved.\n  Traditionally, these are main conference papers, short papers,\n  demos, student research workshop papers, tutorials, and, since 2013,\n  TACL papers.\n- Workshops are single-track: the schedule and all papers are\ncontained in a single session, e.g., `softconf.com/acl2014/WMT14`.\n\nWorkshops are fairly easy to format since everything is linear and\ntypically abstracts are not printed. This can be done fairly\nautomatically by generating the schedule from the order file, even if\nthe workshop chairs are derelict in the their duties.\n\nThe main conference is another matter. The papers that are presented\nare contained in separate subconferences, whose separate schedules must be\nassembled and interleaved to produce the consolidated main conference schedule.\n\n# Notes for Publications Chairs\n\nThe downside of being the handbook chair is that there is all sorts of\nupstream events you have no power to enforce (except through friendly\npub chairs) but that have a large impact on how difficult your job\nis. It will be helpful to read this section and pass this information\non to the publications chairs, in the case where you are not them.\n\nThe single-most important thing you can do is to ensure that the\n\"order\" files are properly formatted, for the workshops, but most\nimportantly, for all parts of the main conference proceedings.\n\nThe \"order\" file is a human- and machine-readable file that is part of\nthe ACLPUB package (which Softconf's START system helps you\nassemble). It is used to generate the schedule in the proceedings,\nwhere computer-readability isn't very important, but it is also used\nto generate the schedules in the handbook, where computer-readability\nis very important, since the schedules from many different workshops\nand pieces of the main conference (papers, shortpapers, SRW, demos,\ntutorials, and, since 2013, TACL) are assembled.\n\nPublications chairs should ensure the following:\n\n- If at all possible, convince the PC chairs **not** to use START's\n  \"ScheduleMaker\" tool. This is a complicated Excel front-end setup\n  that generates the \"order\" file, which is already fairly simple to\n  edit. They should just edit the order file directly. It is much\n  simpler. \n\n- The order files should be machine readable. Workshop chairs\n  are sometimes tempted to play with the formatting lines in\n  order to get a custom look, but this introduces a host of problems\n  generating the handbook.\n  \n  To help ensure compliance, a script has been provided:\n  \n     cat papers/order | ./scripts/verify_schedule.py\n\n- The main conference schedule --- the one containing all sorts of\n  non paper-related events like \"Lunch\" and \"Coffee Break\" and\n  \"Keynote address\" --- should be placed in the \"papers\"\n  subconference, which represents the main conference proceedings. The\n  other order files should contain _only_ events relevant to those\n  workshops. For example, do not repeat coffee breaks and lunches\n  in the \"shortpapers\"; \"shortpapers\" should contain only lines\n  listing days, session titles, and the papers that are presented in\n  those sessions. This is to aid in merging the schedules for these\n  subconferences into a single unified schedule: the code in\n  `scripts/` will group papers in identically-named sessions across\n  files. While events like lunches could also be merged, it is better\n  to just list them in one place so that changes to one file (e.g.,\n  you decide to rename \"Coffee Break\" to \"Tea time\" for this year's\n  meeting in Cambridge) do not require changes to be made all over the\n  place. \n\n- Parallel sessions should be named using the convention \"Session NL: NAME\", where\n  N is a session number (1..as many as necessary), L is a letter (A..number of parallel tracks),\n  and NAME is the session name. Session names should be globally distinct. For example,\n\n       = Session 6D: Machine Translation II\n\n  This aids immensely in generating the handbook and also the poster placards that sit\n  outside the rooms and provide a summary of all papers in each track. \n  \n- If papers are mixed across workshops (e.g., a TACL paper presented\n  with main conference long papers), the session headers should be\n  identical across the order files. e.g., in `papers/order`, you could\n  have\n  \n       = Session 1D: Syntax, Parsing, and Tagging I \n       247 11:00--11:25 # Tagging The Web: Building A Robust Web Tagger with Neural Network   \n\n  and in `tacl/order`, you would put\n  \n       = Session 1D: Syntax, Parsing, and Tagging I\n       6 10:10--10:35 # A Tabular Method for Dynamic Oracles in Transition-Based Parsing\n       5 10:35--11:00 # Joint Incremental Disfluency Detection and Dependency Parsin\n       12 11:25--11:50 # A Crossing-Sensitive Third-Order Factorization for Dependency Parsing\n\n  The identical session name allows them to be merged and formatted\n  automatically when you generate the handbook.\n  \n- TACL papers are a recent addition to *ACL conferences. To facilitate\n  construction of the handbook, the publications chairs should create a dummy\n  subconference named \"TACL\", manually enter the papers and metadata\n  for each of them, and then generate the proceedings from the ACLPUB\n  tab. These will not be published as part of the proceedings, but are\n  used only for the handbook.\n\n# The order file\n\nThe script `scripts/verify_schedule.py` provides some minimal\nsanity-checking on the order file. Here is more detail on how to\nproduce it and how to handle some common limitations.\n\nThe order file permits four types of lines\n\n- Comments\n- Papers. Papers are either presented in oral sessions or in poster\n  sessions. They are denoted by a number, which is tied to the paper\n  submission ID, and is used to locate the paper and the paper's\n  metadata in `proceedings/final/NUM/NUM_metadata.txt`. Papers\n  presented as part of oral sessions have time ranges, while papers\n  presented as posters do not.\n- '+' events, which are single events that have a time range but that\n  are not associated with a paper. For example, lunch.\n- Session headers, denoted by lines starting with '=', which are used\n  to group together papers presented in a parallel track.\n\nHere is an example schedule (from `data/papers/proceedings/order`):\n\n    + 7:30--18:00 Registration\n    + 7:30--9:00 Breakfast\n    + 9:00--10:00 Invited talk II: Zoran Popovic\n    + 10:00--10:30 Coffee break\n    = Session 4A: Machine Learning for NLP\n    321 10:30--10:55 # Kneser-Ney Smoothing on Expected Counts\n    605 10:55--11:20 # Robust Entity Clustering via Phylogenetic Inference\n    377 11:20--11:45 # Linguistic Structured Sparsity in Text Categorization\n    510 11:45--12:10 # Perplexity on Reduced Corpora\n\nIf you want to mix papers between sessions, make sure the session\ntitles are the same. Here are snippets from\n`data/papers/proceedings/order` \n\n    = Session 1D: Syntax, Parsing, and Tagging I\n    247 11:00--11:25 # Tagging The Web: Building A Robust Web Tagger with Neural Network\n\nand `data/tacl/proceedings/order`:\n\n    = Session 1D: Syntax, Parsing, and Tagging I\n    6 10:10--10:35 # A Tabular Method for Dynamic Oracles in Transition-Based Parsing\n    5 10:35--11:00 # Joint Incremental Disfluency Detection and Dependency Parsin\n    12 11:25--11:50 # A Crossing-Sensitive Third-Order Factorization for Dependency Parsing\n\nThese will be assembled automatically. Here's how to do a poster\nsession:\n\n    + 18:50--21:30 Poster and Dinner Session I: TACL Papers, Long Papers, Short Papers, Student Research Workshop; Demonstrations\n    249  # Interpretable Semantic Vectors from a Joint Model of Brain- and Text- Based Meaning\n    416  # Single-Agent vs. Multi-Agent Techniques for Concurrent Reinforcement Learning of Negotiation Dialogue Policies\n    161  # A Linear-Time Bottom-Up Discourse Parser with Constraints and Post-Editing\n    178  # Negation Focus Identification with Contextual Discourse Information\n    ...  \n\n## Missing features\n\nIn START, the code used to verify the order file is not suited to the\nhandbook. Their verification ensures only that each paper in the final\nversion of the proceedings is listed exactly once. This prevents a\nnumber of common situations that people require.\n\n1. Listing a paper more than once in the schedule.\n\n   Some people want this, say, when a paper is provided with both an\n   oral presentation and is also present in a poster session. You can\n   list the paper more than once outside of START, however, and the\n   handbook code will handle it fine. However, you'll want to ensure\n   that the abstract is not printed twice. You'll have to do this by hand.\n\n2. Publishing a paper in the proceedings but not listing it in the\nschedule.\n\n   START won't allow this, but you can just comment out the paper in\n   the schedule for the handbook and all will be well.\n   \n3. Listing a paper in the schedule that is not present in the\nproceedings. \n\n   Sometimes there are papers not in the proceedings, but in the\n   schedule, and you want the formatting to look the same. To\n   accomplish this, you have to manually create new numbered entries\n   for those papers in a separate proceedings tarball. Create new\n   numbered entries and then the corresponding metadata in \n   \n       SUBCONF/proceedings/final/NUM/NUM_metadata.txt\n     \n   The handbook code will then find what it needs and all shall be well.\n\n# Layout\n\nNow we get to assembling the handbook.\n\nACL 2014 had five parallel sessions. This should be useful for any handbook \nwhere parallel sessions are listed one per page with abstracts.\n\nDirectories:\n\n- `input/`: fixed inputs\n   - `input/conferences.txt`: list of conferences (used for auto-download)\n\n- `data/`: where the ACLPUB tarballs are downloaded and unpacked to\n\n- `scripts/`: scripts used to generate the handbook sections from the ACLPUB\n  proceedings.\n\n- `content/`: handbook tex files used to build the handbook\n\n- `auto/`: the output of scripts, used to generate first-pass LaTeX\n  that is then pulled manually into the handbook content.\n\n# Task list\n\n- Download all the main conference proceedings and workshops using\n  `scripts/download-proceedings.sh` (after editing\n  `input/conferences.txt`). This creates a proceedings tarball in each\n  of `data/SUBCONF/proceedings`.\n\n- Verify each of them with `scripts/verify_schedule.py`:\n\n        for file in $(ls data); do \n          num=$(cat data/$file/proceedings/order | ./scripts/verify_schedule.py \u003e /dev/null 2\u003e\u00261; echo $?) \n          if test $num -ne 0; then \n            echo -e \"$file\\t$num\"\n          fi\n        done\n\n- Generate the bibtex metadata from each workshop's paper metadata:\n\n        for dir in $(ls data); do \n          [[ ! -d \"auto/$dir\" ]] \u0026\u0026 mkdir auto/$dir\n          ./scripts/meta2bibtex.py data/$dir/proceedings/final $dir\n        done\n    \n  This creates abstracts in `auto/abstracts` (read in via LaTeX calls\n  to `\\paperabstract`) and BibTeX metadata used for the index and\n  other things (in `auto/SUBCONF/papers.bib`)\n\n- Generate the workshop schedules:\n\n       for dir in $(ls data); do \n         [[ ! -d \"auto/$dir\" ]] \u0026\u0026 mkdir auto/$dir\n         cat data/$dir/proceedings/order | ./scripts/order2schedule_workshop.pl $dir \u003e auto/$dir/schedule.tex\n       done\n    \n  This leaves you with a ton of `schedule.tex` files which can be\n  `\\input`ed via LaTeX\n\n- Generate the paper and poster session files (which you'll have to edit a bit afterwards):\n\n       for name in tacl demos papers; do\n           cat data/$name/proceedings/order | ./scripts/order2schedule.perl $name\n       done\n\n- Edit `content/workshops/overview.tex` and\n  `content/workshops/workshops.tex` to include these files and to be correct.\n\n- Next, fill in the tutorials manually, editing\n  `content/sunday/tutorials-001.tex` and so on. Also edit the tutorial\n  overview page in `content/sunday/sunday.tex`.\n\n- Generate the daily overviews, munge them a bit, pull them in\n\n        cat data/{papers,shortpapers,demos,tacl,srw}/proceedings/order | ./scripts/order2schedule_overview.py\n\n- Email Dragomir Radev, who will run your index against\n  [the ACL Anthology Network](http://clair.eecs.umich.edu/aan/index.php),\n  correcting spellings and collapsing redundancies.\n\n\n# Miscellaneous notes\n\n- Take care to ensure the index is correct. Email Drago Radev who can\n  help you consolidate names against the ACL Anthology.\n\n- Mausam will cause you trouble. Grep for him. You want just the name,\n  no {}s\n  \n- Also special characters in abstracts (e.g., a real alpha, funny\n  latex, chinese, etc). Really this should all be converted over to XeTeX.\n  \n\n# Suggestions for the future\n\nThe ACL proceedings and handbook play complementary roles. In one\nsense, the proceedings are the main product of a conference; their\nlife is a long one, stretching into the future, serving up papers from\nthe Anthology for years to come. However, if that were the only\npurpose of a conference, it would be a journal. The handbook is what\nallows people to quickly and easily navigate the physical space of the\nconference. Its utility may be rooted in a particular time and place,\nbut it is equally a product of the conference and plays an important\nif ephemeral role in bringing people together.\n\nPerhaps in part because of this rooting in a physical space, the task\nof generating the handbook in past years has fallen under the purview\nof the local arrangements chair. This is a mistake. The job is much\nbetter suited to the publications chair. This is true conceptually,\nbut also because many of the tasks of handbook generation (in\nparticular, ensuring machine-readable format of the order files) are\nredundant with work already undergone by the publications chairs, since\na version of the schedule is also published in the proceedings.\n\n# Credits\n\nThis document was written by Matt Post during assembly of the NAACL\n2013 and ACL 2014 handbooks. I inherited from Ulrich Germann the code\nand data he used to assemble the 2012 NAACL handbook. I don't know\nabout any history prior to that.\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmjpost%2Facl2014-handbook","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmjpost%2Facl2014-handbook","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmjpost%2Facl2014-handbook/lists"}