Ecosyste.ms: Awesome
An open API service indexing awesome lists of open source software.
https://github.com/apache/orc
Apache ORC - the smallest, fastest columnar storage for Hadoop workloads
https://github.com/apache/orc
apache big-data cpp java orc
Last synced: 3 days ago
JSON representation
Apache ORC - the smallest, fastest columnar storage for Hadoop workloads
- Host: GitHub
- URL: https://github.com/apache/orc
- Owner: apache
- License: apache-2.0
- Created: 2015-05-06T07:00:05.000Z (over 9 years ago)
- Default Branch: main
- Last Pushed: 2024-12-30T12:38:48.000Z (12 days ago)
- Last Synced: 2025-01-02T00:06:33.296Z (10 days ago)
- Topics: apache, big-data, cpp, java, orc
- Language: Java
- Homepage: https://orc.apache.org/
- Size: 54.1 MB
- Stars: 703
- Watchers: 48
- Forks: 485
- Open Issues: 31
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
- stars - apache/orc - the smallest, fastest columnar storage for Hadoop workloads (HarmonyOS / Windows Manager)
- awesome-dataops - Apache ORC - A self-describing type-aware columnar file format designed for Hadoop workloads. (Data Serialization)
- awesome-datalake - Apache ORC - ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. (File Formats)
- awesome-datalake - Apache ORC - ORC is a self-describing type-aware columnar file format designed for Hadoop workloads. (File Formats)
README
# [Apache ORC](https://orc.apache.org/)
ORC is a self-describing type-aware columnar file format designed for
Hadoop workloads. It is optimized for large streaming reads, but with
integrated support for finding required rows quickly. Storing data in
a columnar format lets the reader read, decompress, and process only
the values that are required for the current query. Because ORC files
are type-aware, the writer chooses the most appropriate encoding for
the type and builds an internal index as the file is written.
Predicate pushdown uses those indexes to determine which stripes in a
file need to be read for a particular query and the row indexes can
narrow the search to a particular set of 10,000 rows. ORC supports the
complete set of types in Hive, including the complex types: structs,
lists, maps, and unions.## ORC File Library
This project includes both a Java library and a C++ library for reading and writing the _Optimized Row Columnar_ (ORC) file format. The C++ and Java libraries are completely independent of each other and will each read all versions of ORC files.
Releases:
* Latest: Apache ORC releases
* Maven Central: ![Maven Central](https://maven-badges.herokuapp.com/maven-central/org.apache.orc/orc/badge.svg)
* Downloads: Apache ORC downloads
* Release tags: Apache ORC release tags
* Plan: Apache ORC future release planThe current build status:
* Main branch
![main build status](https://github.com/apache/orc/actions/workflows/build_and_test.yml/badge.svg?branch=main)Bug tracking: Apache Jira
The subdirectories are:
* c++ - the c++ reader and writer
* cmake_modules - the cmake modules
* docker - docker scripts to build and test on various linuxes
* examples - various ORC example files that are used to test compatibility
* java - the java reader and writer
* site - the website and documentation
* tools - the c++ tools for reading and inspecting ORC files### Building
* Install java 17 or higher
* Install maven 3.9.9 or higher
* Install cmake 3.12 or higherTo build a release version with debug information:
```shell
% mkdir build
% cd build
% cmake ..
% make package
% make test-out```
To build a debug version:
```shell
% mkdir build
% cd build
% cmake .. -DCMAKE_BUILD_TYPE=DEBUG
% make package
% make test-out```
To build a release version without debug information:
```shell
% mkdir build
% cd build
% cmake .. -DCMAKE_BUILD_TYPE=RELEASE
% make package
% make test-out```
To build only the Java library:
```shell
% cd java
% ./mvnw package```
To build only the C++ library:
```shell
% mkdir build
% cd build
% cmake .. -DBUILD_JAVA=OFF
% make package
% make test-out```
To build the C++ library with AVX512 enabled:
```shell
export ORC_USER_SIMD_LEVEL=AVX512
% mkdir build
% cd build
% cmake .. -DBUILD_JAVA=OFF -DBUILD_ENABLE_AVX512=ON
% make package
% make test-out
```
Cmake option BUILD_ENABLE_AVX512 can be set to "ON" or (default value)"OFF" at the compile time. At compile time, it defines the SIMD level(AVX512) to be compiled into the binaries.Environment variable ORC_USER_SIMD_LEVEL can be set to "AVX512" or (default value)"NONE" at the run time. At run time, it defines the SIMD level to dispatch the code which can apply SIMD optimization.
Note that if ORC_USER_SIMD_LEVEL is set to "NONE" at run time, AVX512 will not take effect at run time even if BUILD_ENABLE_AVX512 is set to "ON" at compile time.