{"id":38000899,"url":"https://github.com/fraugster/parquet-go","last_synced_at":"2026-01-16T19:12:52.697Z","repository":{"id":38452171,"uuid":"257618953","full_name":"fraugster/parquet-go","owner":"fraugster","description":"Go package to read and write parquet files. parquet is a file format to store nested data structures in a flat columnar data format. It can be used in the Hadoop ecosystem and with tools such as Presto and AWS Athena.","archived":false,"fork":false,"pushed_at":"2024-04-09T16:11:03.000Z","size":1254,"stargazers_count":279,"open_issues_count":19,"forks_count":53,"subscribers_count":11,"default_branch":"master","last_synced_at":"2024-06-18T12:40:27.140Z","etag":null,"topics":["athena","golang","golang-package","hacktoberfest","hadoop","parquet","parquet-schema","presto"],"latest_commit_sha":null,"homepage":"","language":"Go","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/fraugster.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2020-04-21T14:17:49.000Z","updated_at":"2024-06-17T23:39:33.000Z","dependencies_parsed_at":"2024-06-18T12:29:33.171Z","dependency_job_id":"d9288da4-c637-4e87-8418-3b549515dc49","html_url":"https://github.com/fraugster/parquet-go","commit_stats":null,"previous_names":[],"tags_count":15,"template":false,"template_full_name":null,"purl":"pkg:github/fraugster/parquet-go","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fraugster%2Fparquet-go","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fraugster%2Fparquet-go/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fraugster%2Fparquet-go/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fraugster%2Fparquet-go/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/fraugster","download_url":"https://codeload.github.com/fraugster/parquet-go/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/fraugster%2Fparquet-go/sbom","scorecard":{"id":409999,"data":{"date":"2025-08-11","repo":{"name":"github.com/fraugster/parquet-go","commit":"0c50c9da7cd7835c30640c7f51401547bf79d30f"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":3.2,"checks":[{"name":"Token-Permissions","score":-1,"reason":"No tokens found","details":null,"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Code-Review","score":3,"reason":"Found 6/16 approved changesets -- score normalized to 3","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Dangerous-Workflow","score":-1,"reason":"no workflows found","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: Apache License 2.0: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":-1,"reason":"internal error: error during branchesHandler.setup: internal error: githubv4.Query: Resource not accessible by integration","details":null,"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: containerImage not pinned by hash: compatibility/Dockerfile:1","Warn: containerImage not pinned by hash: compatibility/Dockerfile:15","Warn: containerImage not pinned by hash: compatibility/Dockerfile:24: pin your Docker image by updating openjdk:8-jdk-alpine to openjdk:8-jdk-alpine@sha256:94792824df2df33402f201713f932b58cb9de94a0cd524164a0f2283343547b3","Info:   0 out of   3 containerImage dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 20 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}},{"name":"Vulnerabilities","score":7,"reason":"3 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: GO-2021-0061 / GHSA-r88r-gmrh-7j83","Warn: Project is vulnerable to: GO-2022-0956 / GHSA-6q6q-88xp-6f2r","Warn: Project is vulnerable to: GO-2020-0036 / GHSA-wxc4-f4m6-wwqv"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}}]},"last_synced_at":"2025-08-18T22:26:11.239Z","repository_id":38452171,"created_at":"2025-08-18T22:26:11.239Z","updated_at":"2025-08-18T22:26:11.239Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28481550,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-16T11:59:17.896Z","status":"ssl_error","status_checked_at":"2026-01-16T11:55:55.838Z","response_time":107,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["athena","golang","golang-package","hacktoberfest","hadoop","parquet","parquet-schema","presto"],"created_at":"2026-01-16T19:12:52.549Z","updated_at":"2026-01-16T19:12:52.691Z","avatar_url":"https://github.com/fraugster.png","language":"Go","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003ch1 align=\"center\"\u003eparquet-go\u003c/h1\u003e\n\u003cp align=\"center\"\u003e\n        \u003ca href=\"https://github.com/fraugster/parquet-go/releases\"\u003e\u003cimg src=\"https://img.shields.io/github/v/tag/fraugster/parquet-go.svg?color=brightgreen\u0026label=version\u0026sort=semver\"\u003e\u003c/a\u003e\n        \u003ca href=\"https://circleci.com/gh/fraugster/parquet-go/tree/master\"\u003e\u003cimg src=\"https://circleci.com/gh/fraugster/parquet-go/tree/master.svg?style=shield\"\u003e\u003c/a\u003e\n        \u003ca href=\"https://goreportcard.com/report/github.com/fraugster/parquet-go\"\u003e\u003cimg src=\"https://goreportcard.com/badge/github.com/fraugster/parquet-go\"\u003e\u003c/a\u003e\n        \u003ca href=\"https://codecov.io/gh/fraugster/parquet-go\"\u003e\u003cimg src=\"https://codecov.io/gh/fraugster/parquet-go/branch/master/graph/badge.svg\"/\u003e\u003c/a\u003e\n        \u003ca href=\"https://godoc.org/github.com/fraugster/parquet-go\"\u003e\u003cimg src=\"https://img.shields.io/badge/godoc-reference-blue.svg?color=blue\"\u003e\u003c/a\u003e\n        \u003ca href=\"https://github.com/fraugster/parquet-go/blob/master/LICENSE\"\u003e\u003cimg src=\"https://img.shields.io/badge/license-Apache%202-blue\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n---\n\nparquet-go is an implementation of the [Apache Parquet file format](https://github.com/apache/parquet-format)\nin Go. It provides functionality to both read and write parquet files, as well\nas high-level functionality to manage the data schema of parquet files, to\ndirectly write Go objects to parquet files using automatic or custom\nmarshalling and to read records from parquet files into Go objects using\nautomatic or custom marshalling.\n\nparquet is a file format to store nested data structures in a flat columnar\nformat. By storing in a column-oriented way, it allows for efficient reading\nof individual columns without having to read and decode complete rows. This\nallows for efficient reading and faster processing when using the file format\nin conjunction with distributed data processing frameworks like Apache Hadoop\nor distributed SQL query engines like Presto and AWS Athena.\n\nThis implementation is divided into several packages. The top-level package is\nthe low-level implementation of the parquet file format. It is accompanied by\nthe sub-packages parquetschema and floor. parquetschema provides functionality\nto parse textual schema definitions as well as the data types to manually or\nprogrammatically construct schema definitions. floor is a high-level wrapper\naround the low-level package. It provides functionality to open parquet files\nto read from them or write to them using automated or custom marshalling and\nunmarshalling.\n\n## Supported Features\n\n| Feature                                  | Read | Write | Note |\n| ---                                      | ---- | ---- | --- |\n| Compression                              | Yes  | Yes  | Only GZIP and SNAPPY are supported out of the box, but it is possible to add other compressors, see below. |\n| Dictionary Encoding                      | Yes  | Yes  |\n| Run Length Encoding / Bit-Packing Hybrid | Yes  | Yes  | The reader can read RLE/Bit-pack encoding, but the writer only uses bit-packing |\n| Delta Encoding                           | Yes  | Yes  |\n| Byte Stream Split                        | No   | No   |\n| Data page V1                             | Yes  | Yes  |\n| Data page V2                             | Yes  | Yes  |\n| Statistics in page meta data             | No   | Yes  | Page meta data is generally not made available to users and not used by parquet-go.\n| Index Pages                              | No   | No   |\n| Dictionary Pages                         | Yes  | Yes  |\n| Encryption                               | No   | No   |\n| Bloom Filter                             | No   | No   |\n| Logical Types                            | Yes  | Yes  | Support for logical type is in the high-level package (floor) the low level parquet library only supports the basic types, see the type mapping table |\n\n## Supported Data Types\n\n| Type in parquet         | Type in Go      | Note |\n| ----------------------- | --------------- | ---- |\n| boolean                 | bool            |\n| int32                   | int32           | See the note about the int type |\n| int64                   | int64           | See the note about the int type |\n| int96                   | [12]byte        |\n| float                   | float32         |\n| double                  | float64         |\n| byte_array              | []byte          |\n| fixed_len_byte_array(N) | [N]byte, []byte | use any positive number for `N` |\n\nNote: the low-level implementation only supports int32 for the INT32 type and int64 for the INT64 type in Parquet.\nPlain int or uint are not supported. The high-level `floor` package contains more extensive support for these\ndata types.\n\n## Supported Logical Types\n\n| Logical Type   | Mapped to Go types      | Note |\n| -------------- | ----------------------- | ---- |\n| STRING         | string, []byte          |\n| DATE           | int32, time.Time        | int32: days since Unix epoch (Jan 01 1970 00:00:00 UTC); time.Time only in `floor` |\n| TIME           | int32, int64, time.Time | int32: TIME(MILLIS, ...), int64: TIME(MICROS, ...), TIME(NANOS, ...); time.Time only in `floor` |\n| TIMESTAMP      | int64, int96, time.Time | time.Time only in `floor`\n| UUID           | [16]byte                |\n| LIST           | []T                     | slices of any type |\n| MAP            | map[T1]T2               | maps with any key and value types |\n| ENUM           | string, []byte          |\n| BSON           | []byte                  |\n| DECIMAL        | []byte, [N]byte         |\n| INT            | {,u}int{8,16,32,64}     | implementation is loose and will allow any INT logical type converted to any signed or unsigned int Go type. |\n\n## Supported Converted Types\n\n| Converted Type       | Mapped to Go types  | Note |\n| -------------------- | ------------------- | ---- |\n| UTF8                 | string, []byte      |\n| TIME\\_MILLIS          | int32               | Number of milliseconds since the beginning of the day |\n| TIME\\_MICROS          | int64               | Number of microseconds since the beginning of the day |\n| TIMESTAMP\\_MILLIS     | int64               | Number of milliseconds since Unix epoch (Jan 01 1970 00:00:00 UTC) |\n| TIMESTAMP\\_MICROS     | int64               | Number of milliseconds since Unix epoch (Jan 01 1970 00:00:00 UTC) |\n| {,U}INT\\_{8,16,32,64} | {,u}int{8,16,32,64} | implementation is loose and will allow any converted type with any int Go type. |\n| INTERVAL             | [12]byte            |\n\nPlease note that converted types are deprecated. Logical types should be used preferably.\n\n## Supported Compression Algorithms\n\n| Compression Algorithm | Supported | Notes |\n| --------------------- | --------- | ----- |\n| GZIP                  | Yes; Out of the box |\n| SNAPPY                | Yes; Out of the box |\n| BROTLI                | Yes; By importing [github.com/akrennmair/parquet-go-brotli](https://github.com/akrennmair/parquet-go-brotli) |\n| LZ4                   | No | LZ4 has been deprecated as of parquet-format 2.9.0. |\n| LZ4\\_RAW              | Yes; By importing [github.com/akrennmair/parquet-go-lz4raw](https://github.com/akrennmair/parquet-go-lz4raw) |\n| LZO                   | Yes; By importing [github.com/akrennmair/parquet-go-lzo](https://github.com/akrennmair/parquet-go-lzo) | Uses a cgo wrapper around the original LZO implementation which is licensed as GPLv2+. |\n| ZSTD                  | Yes; By importing [github.com/akrennmair/parquet-go-zstd](https://github.com/akrennmair/parquet-go-zstd) |\n\n## Schema Definition\n\nparquet-go comes with support for textual schema definitions. The sub-package\n`parquetschema` comes with a parser to turn the textual schema definition into\nthe right data type to use elsewhere to specify parquet schemas. The syntax\nhas been mostly reverse-engineered from a similar format also supported but\nbarely documented in [Parquet's Java implementation](https://github.com/apache/parquet-mr/blob/master/parquet-column/src/main/java/org/apache/parquet/schema/MessageTypeParser.java).\n\nFor the full syntax, please have a look at the [parquetschema package Go documentation](http://godoc.org/github.com/fraugster/parquet-go/parquetschema).\n\nGenerally, the schema definition describes the structure of a message. Parquet\nwill then flatten this into a purely column-based structure when writing the\nactual data to parquet files.\n\nA message consists of a number of fields. Each field either has type or is a\ngroup. A group itself consists of a number of fields, which in turn can have\neither a type or are a group themselves. This allows for theoretically\nunlimited levels of hierarchy.\n\nEach field has a repetition type, describing whether a field is required (i.e.\na value has to be present), optional (i.e. a value can be present but doesn't\nhave to be) or repeated (i.e. zero or more values can be present). Optionally,\neach field (including groups) have an annotation, which contains a logical type\nor converted type that annotates something about the general structure at this\npoint, e.g. `LIST` indicates a more complex list structure, or `MAP` a key-value\nmap structure, both following certain conventions. Optionally, a typed field\ncan also have a numeric field ID. The field ID has no purpose intrinsic to the\nparquet file format.\n\nHere is a simple example of a message with a few typed fields:\n\n```\nmessage coordinates {\n    required float64 latitude;\n    required float64 longitude;\n    optional int32 elevation = 1;\n    optional binary comment (STRING);\n}\n```\n\nIn this example, we have a message with four typed fields, two of them\nrequired, and two of them optional. `float64`, `int32` and `binary` describe\nthe fundamental data type of the field, while `longitude`, `latitude`,\n`elevation` and `comment` are the field names. The parentheses contain\nan annotation `STRING` which indicates that the field is a string, encoded\nas binary data, i.e. a byte array. The field `elevation` also has a field\nID of `1`, indicated as numeric literal and separated from the field name\nby the equal sign `=`.\n\nIn the following example, we will introduce a plain group as well as two\nnested groups annotated with logical types to indicate certain data structures:\n\n```\nmessage transaction {\n    required fixed_len_byte_array(16) txn_id (UUID);\n    required int32 amount;\n    required int96 txn_ts;\n    optional group attributes {\n        optional int64 shop_id;\n        optional binary country_code (STRING);\n        optional binary postcode (STRING);\n    }\n    required group items (LIST) {\n        repeated group list {\n            required int64 item_id;\n            optional binary name (STRING);\n        }\n    }\n    optional group user_attributes (MAP) {\n        repeated group key_value {\n            required binary key (STRING);\n            required binary value (STRING);\n        }\n    }\n}\n```\n\nIn this example, we see a number of top-level fields, some of which are\ngroups. The first group is simply a group of typed fields, named `attributes`.\n\nThe second group, `items` is annotated to be a `LIST` and in turn contains a\n`repeated group list`, which in turn contains a number of typed fields. When\na group is annotated as `LIST`, it needs to follow a particular convention:\nit has to contain a `repeated group` named `list`. Inside this group, any\nfields can be present.\n\nThe third group, `user_attributes` is annotated as `MAP`. Similar to `LIST`,\nit follows some conventions. In particular, it has to contain only a single\n`required group` with the name `key_value`, which in turn contains exactly two\nfields, one named `key`, the other named `value`. This represents a map\nstructure in which each key is associated with one value.\n\n## Examples\n\nFor examples how to use both the low-level and high-level APIs of this library, please\nsee the directory `examples`. You can also check out the accompanying tools (see below)\nfor more advanced examples. The tools are located in the `cmd` directory.\n\n## Tools\n\n`parquet-go` comes with tooling to inspect and generate parquet tools.\n\n### parquet-tool\n\n`parquet-tool` allows you to inspect the meta data, the schema and the number of rows\nas well as print the content of a parquet file. You can also use it to split an existing\nparquet file into multiple smaller files.\n\nInstall it by running `go get github.com/fraugster/parquet-go/cmd/parquet-tool` on your command line.\nFor more detailed help on how to use the tool, consult `parquet-tool --help`.\n\n### csv2parquet\n\n`csv2parquet` makes it possible to convert an existing CSV file into a parquet file. By default,\nall columns are simply turned into strings, but you provide it with type hints to influence\nthe generated parquet schema.\n\nYou can install this tool by running `go get github.com/fraugster/parquet-go/cmd/csv2parquet` on your command line.\nFor more help, consult `csv2parquet --help`.\n\n## Contributing\n\nIf you want to hack on this repository, please read the short [CONTRIBUTING.md](CONTRIBUTING.md)\nguide first.\n\n# Versioning\n\nWe use [SemVer](http://semver.org/) for versioning. For the versions available,\nsee the [tags on this repository][tags].\n\n## Authors\n\n- **Forud Ghafouri** - *Initial work* [fzerorubigd](https://github.com/fzerorubigd)\n- **Andreas Krennmair** - *floor package, schema parser* [akrennmair](https://github.com/akrennmair)\n- **Stefan Koshiw** - *Engineering Manager for Core Team* [panamafrancis](https://github.com/panamafrancis)\n\nSee also the list of [contributors][contributors] who participated in this project.\n\n## Special Mentions\n\n- **Nathan Hanna** - *proposal and prototyping of automatic schema generator* [jnathanh](https://github.com/jnathanh)\n\n## License\n\nCopyright 2021 Fraugster GmbH\n\nThis project is licensed under the Apache-2 License - see the [LICENSE](LICENSE) file for details.\n\nThis program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. \n\n[tags]: https://github.com/fraugster/parquet-go/tags\n[contributors]: https://github.com/fraugster/parquet-go/graphs/contributors\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffraugster%2Fparquet-go","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ffraugster%2Fparquet-go","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffraugster%2Fparquet-go/lists"}