{"id":18449434,"url":"https://github.com/public-transport/why-linked-open-transit-data","last_synced_at":"2026-01-24T00:50:37.864Z","repository":{"id":89996767,"uuid":"244886092","full_name":"public-transport/why-linked-open-transit-data","owner":"public-transport","description":"Why do we need linked open public transport data?","archived":false,"fork":false,"pushed_at":"2021-07-23T19:09:27.000Z","size":8,"stargazers_count":21,"open_issues_count":3,"forks_count":1,"subscribers_count":7,"default_branch":"main","last_synced_at":"2025-04-12T23:25:02.597Z","etag":null,"topics":["linked-data","open-data","public-transport","transit"],"latest_commit_sha":null,"homepage":"https://github.com/public-transport/why-linked-open-transit-data#why-linked-open-transit-data","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc0-1.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/public-transport.png","metadata":{"files":{"readme":"readme.md","changelog":null,"contributing":null,"funding":null,"license":"license","code_of_conduct":"code-of-conduct.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2020-03-04T11:44:12.000Z","updated_at":"2025-04-12T20:16:24.000Z","dependencies_parsed_at":"2023-05-30T21:00:23.303Z","dependency_job_id":null,"html_url":"https://github.com/public-transport/why-linked-open-transit-data","commit_stats":{"total_commits":2,"total_committers":1,"mean_commits":2.0,"dds":0.0,"last_synced_commit":"49390ec3126d01ee96d3b2301acd01095c80b2e5"},"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/public-transport%2Fwhy-linked-open-transit-data","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/public-transport%2Fwhy-linked-open-transit-data/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/public-transport%2Fwhy-linked-open-transit-data/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/public-transport%2Fwhy-linked-open-transit-data/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/public-transport","download_url":"https://codeload.github.com/public-transport/why-linked-open-transit-data/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":249117312,"owners_count":21215365,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["linked-data","open-data","public-transport","transit"],"created_at":"2024-11-06T07:20:01.102Z","updated_at":"2026-01-24T00:50:37.836Z","avatar_url":"https://github.com/public-transport.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# Why linked open transit data?\n\nThis repo explains why we need [linked](https://en.wikipedia.org/wiki/Linked_data) open public transport data.\n\n![CC0-licensed](https://img.shields.io/github/license/public-transport/why-linked-open-transit-data.svg)\n[![chat on gitter](https://badges.gitter.im/public-transport/Lobby.svg)](https://gitter.im/public-transport/Lobby)\n\n*Note:* This document is inspired by [*Publishing Transport Data for Maximum Reuse*](https://phd.pietercolpaert.be) by [Pieter Colpaert](https://pietercolpaert.be), which offers a very detailed view on the topic. We recommend to read it!\n\n\n## Problem\n\nWhen travelling through larger regions or several countries by public transportation, finding out how and when to get to the destination is hard:\n\n1. People often **need to use multiple, regionally limited apps** to find out which trains/busses/ferries/etc. are available, because these apps often have imprecise (e.g. regarding accessibility), outdated (e.g. construction work) or just no data whatsoever about other regions. Doing this research across operator boundaries involves a lot of manual work. Essentially the user needs to do the job that computers should do: routing through sub-networks.\n2. When dealing with large distances (e.g. from Norway to France), this routing work becomes almost impossible for humans to do ad-hoc, because there are so many possible connections. Combined with e.g. cancellations \u0026 delays, **users may never find the optimal connection** because of that.\n3. **Local, narrow-focused apps are not (as) accessible.** They're often developed with a smaller budget, in some languages, without screen reader \u0026 offline support, have a bad UX, are only available for some platforms, etc.\n4. **Apps built for current mainstream use cases are not future-proof.** With the ongoing digitisation, diversification and increased on-demand features, they won't be able to deliver on people's mobility needs. (They barely do that *right now*.)\n\n\n## Data Hubs\n\n**An often-proposed (alleged) solution is to build data exchange hubs**: They collect individual data sets (of both plan \u0026 realtime data), integrate them – often using hand-written matching tables and fuzzy matching – and emit one large data merged set. **This doesn't work** for the following reasons:\n\n- **It doesn't scale.** Currently, transportation data hubs process data from *dozens or hundreds* of operators. In Europe, with the on-going heterogenisation of mobility, to really cover *all mobility options*, they would have to integrate *thousands to ten-thousands* of operators.\n- Almost always, **the raw \u0026 merged data is not (easily) accessible**. This makes it hard to a) integrate more data sources (because of missing openly available examples \u0026 APIs) and b) improve the merging algorithms.\n- **It is neither fair nor innovative.** Designing centralized systems (including a federation of central hubs) disproportionately helps large transportation operators and app companies (Google Maps), and actively harms small, innovative operators from appearing in people's day-to-day mobility apps.\n- **It's not resilient.** With the coming level of technical integration in the mobility sector, large centralized data hubs will a) affect a huge number of people when they're down/malfunctioning, b) make it harder to build offline-capable, data-saving, low-end-devices-compatible (IoT) clients, and therefore encourage the reliance on always-on cloud systems.\n\n\n## Linked Open Transport Data\n\nLet's solve these problems by designing our public transportation systems *from the start* **with federation, discovery of data sources and caching/offline compatibility in mind**!\n\nWe must make our data\n\n- **descriptive** (referencing human- \u0026 machine-readable documents on how to parse \u0026 interpret it),\n- **globally precise** (using IDs that accurately identify transportation infrastructure *without local or operator-specific context*),\n- **redistributable** (legally, by putting it in the [public domain](https://en.wikipedia.org/wiki/Public_domain), and technically, by enabling chunking, replication \u0026 caching).\n\nWe must make our APIs\n\n- **openly documented** (using open standards),\n- **publicly available** (without authentication, in the internet, in a best-effort approach),\n- **federated** (linking to other APIs instead of aggregating them).\n\nWe must develop our **data standards** in the open (allowing barrier-free participation \u0026 collaboration), and make them **freely licensed** (to enable wide-spread use). They **should cleanly separate semantics, replication/transport, storage \u0026 encoding**. They should not reinvent the wheel, but **rely on existing work** (such as [GTFS](https://gtfs.org)) where applicable.\n\n\n## Stable Identifiers\n\nBecause public transportation *data* reflects strongly interconnected public transportation *systems*, it has many links. **When data by an author/source \"A\" refers to data from *another* author/source \"B\", it needs a reliable and precise way to identify items in \"B\" data.** In federated systems, especially in [linked data](https://en.wikipedia.org/wiki/Linked_data) systems, the need for stable \u0026 globally unique IDs is even more significant than in traditional, centralized systems.\n\n*Note:* The aforementioned [*Publishing Transport Data for Maximum Reuse*](https://phd.pietercolpaert.be) has a specific [section on stable identifiers for interoperable data](https://phd.pietercolpaert.be/chapters/measuring-iop).\n\n\n---\n\n## Contributing\n\nContributions are welcome! If you have a question or want to propose changes, go to [the Issues page](https://github.com/public-transport/why-linked-open-transit-data/issues). By participating in this project, you commit to the [code of conduct](code-of-conduct.md).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpublic-transport%2Fwhy-linked-open-transit-data","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fpublic-transport%2Fwhy-linked-open-transit-data","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpublic-transport%2Fwhy-linked-open-transit-data/lists"}