Ecosyste.ms: Awesome
An open API service indexing awesome lists of open source software.
https://github.com/Netflix/Surus
https://github.com/Netflix/Surus
Last synced: about 2 months ago
JSON representation
- Host: GitHub
- URL: https://github.com/Netflix/Surus
- Owner: Netflix
- License: apache-2.0
- Created: 2015-01-12T20:31:53.000Z (over 9 years ago)
- Default Branch: master
- Last Pushed: 2023-03-24T09:37:02.000Z (over 1 year ago)
- Last Synced: 2024-07-31T21:53:40.122Z (about 2 months ago)
- Language: Java
- Size: 1.11 MB
- Stars: 458
- Watchers: 491
- Forks: 108
- Open Issues: 16
-
Metadata Files:
- Readme: README.md
- License: LICENSE
Awesome Lists containing this project
README
[![NetflixOSS Lifecycle](https://img.shields.io/osslifecycle/Netflix/Surus.svg)]()
# Surus
A collection of tools for analysis in Pig and Hive.
## Description
Over the next year we plan to release a handful of our internal user defined functions (UDFs) that have broad adoption across Netflix. The use
cases for these functions are varied in nature (e.g. scoring predictive models, outlier detection, pattern matching, etc.) and together extend
the analytical capabilities of big data.## Functions
* ScorePMML - A tool for scoring predictive models in the cloud.
* Robust Anomaly Detection (RAD) - An implementation of the Robust PCA.## Building Surus
Surus is a standard Maven project. After cloning the git repository you can simply run the following command from the project root directory:
mvn clean package
On the first build, Maven will download all the dependencies from the internet and cache them in the local repository (`~/.m2/repository`), which
can take a considerable amount of time. Subsequent builds will be faster.## Using Surus
After building Surus you will need to move it to your Hive/Pig instance and register the JAR in your environment. For those
unfamiliar with this process see the [Apache Pig UDF](https://pig.apache.org/docs/r0.14.0/udf.html),
and [Hive Plugin](https://cwiki.apache.org/confluence/display/Hive/HivePlugins), documentation.You can also install the anomaly detection R package trivially with this code
library(devtools)
install_github(repo = "Surus", username = "Netflix", subdir = "resources/R/RAD")