{"id":13474092,"url":"https://github.com/ray-project/lightgbm_ray","last_synced_at":"2025-10-12T22:43:58.084Z","repository":{"id":37911952,"uuid":"376087328","full_name":"ray-project/lightgbm_ray","owner":"ray-project","description":"LightGBM on Ray","archived":false,"fork":false,"pushed_at":"2024-02-04T19:26:36.000Z","size":169,"stargazers_count":50,"open_issues_count":5,"forks_count":7,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-08-29T16:47:10.475Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ray-project.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2021-06-11T16:45:43.000Z","updated_at":"2025-07-30T10:53:13.000Z","dependencies_parsed_at":"2023-12-28T00:59:36.922Z","dependency_job_id":"b622b4ae-ae69-4b2e-b707-cc17e795544c","html_url":"https://github.com/ray-project/lightgbm_ray","commit_stats":{"total_commits":143,"total_committers":5,"mean_commits":28.6,"dds":0.04195804195804198,"last_synced_commit":"e393d148824d616a3a4fc4afeb1cc7508935b898"},"previous_names":[],"tags_count":13,"template":false,"template_full_name":null,"purl":"pkg:github/ray-project/lightgbm_ray","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ray-project%2Flightgbm_ray","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ray-project%2Flightgbm_ray/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ray-project%2Flightgbm_ray/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ray-project%2Flightgbm_ray/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ray-project","download_url":"https://codeload.github.com/ray-project/lightgbm_ray/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ray-project%2Flightgbm_ray/sbom","scorecard":{"id":763367,"data":{"date":"2025-08-11","repo":{"name":"github.com/ray-project/lightgbm_ray","commit":"4c4d3413f86db769bddb6d08e2480a04bc75d712"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":3.2,"checks":[{"name":"Code-Review","score":5,"reason":"Found 15/30 approved changesets -- score normalized to 5","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Dangerous-Workflow","score":10,"reason":"no dangerous workflow patterns detected","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Token-Permissions","score":0,"reason":"detected GitHub workflow tokens with excessive permissions","details":["Warn: no topLevel permission defined: .github/workflows/gpu.yaml:1","Warn: no topLevel permission defined: .github/workflows/test.yaml:1","Info: no jobLevel write permissions found"],"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/gpu.yaml:11: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/gpu.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/gpu.yaml:13: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/gpu.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:82: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:84: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:99: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:105: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:120: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:122: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:142: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:148: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:169: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:171: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:186: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:203: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:209: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:14: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:16: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:46: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yaml:48: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:63: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test.yaml:69: update your workflow using https://app.stepsecurity.io/secureworkflow/ray-project/lightgbm_ray/test.yaml/main?enable=pin","Warn: pipCommand not pinned by hash: lightgbm_ray/tests/release/setup_lightgbm.sh:3","Warn: pipCommand not pinned by hash: lightgbm_ray/tests/release/setup_lightgbm.sh:8","Warn: pipCommand not pinned by hash: .github/workflows/gpu.yaml:18","Warn: pipCommand not pinned by hash: .github/workflows/gpu.yaml:19","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:21","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:22","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:23","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:53","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:54","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:55","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:58","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:89","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:90","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:91","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:94","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:127","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:128","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:129","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:137","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:176","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:177","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:178","Warn: pipCommand not pinned by hash: .github/workflows/test.yaml:184","Info:   0 out of  13 GitHub-owned GitHubAction dependencies pinned","Info:   0 out of   8 third-party GitHubAction dependencies pinned","Info:   0 out of  23 pipCommand dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: Apache License 2.0: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":0,"reason":"Project has not signed or included provenance with any releases.","details":["Warn: release artifact v0.1.8 not signed: https://api.github.com/repos/ray-project/lightgbm_ray/releases/81706437","Warn: release artifact v0.1.5 not signed: https://api.github.com/repos/ray-project/lightgbm_ray/releases/74530294","Warn: release artifact v0.1.8 does not have provenance: https://api.github.com/repos/ray-project/lightgbm_ray/releases/81706437","Warn: release artifact v0.1.5 does not have provenance: https://api.github.com/repos/ray-project/lightgbm_ray/releases/74530294"],"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":-1,"reason":"internal error: error during branchesHandler.setup: internal error: githubv4.Query: Resource not accessible by integration","details":null,"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"Vulnerabilities","score":2,"reason":"8 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: PYSEC-2024-48 / GHSA-fj7x-q9j7-g6q6","Warn: Project is vulnerable to: GHSA-3pww-qvr8-6mhp","Warn: Project is vulnerable to: GHSA-6cxr-8q3m-jwrr","Warn: Project is vulnerable to: GHSA-h3xg-wv58-5p43","Warn: Project is vulnerable to: PYSEC-2025-23 / GHSA-w4rh-fgx7-q63m","Warn: Project is vulnerable to: PYSEC-2020-107 / GHSA-jjw5-xxj6-pcv5","Warn: Project is vulnerable to: PYSEC-2024-110 / GHSA-jw8x-6495-233v","Warn: Project is vulnerable to: PYSEC-2020-108"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 17 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}}]},"last_synced_at":"2025-08-23T00:03:02.002Z","repository_id":37911952,"created_at":"2025-08-23T00:03:02.003Z","updated_at":"2025-08-23T00:03:02.003Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279013281,"owners_count":26085250,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-12T02:00:06.719Z","response_time":53,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-31T16:01:09.466Z","updated_at":"2025-10-12T22:43:58.061Z","avatar_url":"https://github.com/ray-project.png","language":"Python","funding_links":[],"categories":["Python","Models and Projects"],"sub_categories":["Ray + X (integration)"],"readme":"\u003c!--$UNCOMMENT(lightgbm-ray)=--\u003e\n\n# Distributed LightGBM on Ray\n\u003c!--$REMOVE--\u003e\n![Build Status](https://github.com/ray-project/lightgbm_ray/workflows/pytest%20on%20push/badge.svg)\n[![docs.ray.io](https://img.shields.io/badge/docs-ray.io-blue)](https://docs.ray.io/en/master/lightgbm-ray.html)\n\u003c!--$END_REMOVE--\u003e\nLightGBM-Ray is a distributed backend for\n[LightGBM](https://lightgbm.readthedocs.io/), built\non top of\n[distributed computing framework Ray](https://ray.io).\n\nLightGBM-Ray\n\n- enables [multi-node](#usage) and [multi-GPU](#multi-gpu-training) training\n- integrates seamlessly with distributed [hyperparameter optimization](#hyperparameter-tuning) library [Ray Tune](http://tune.io)\n- comes with [fault tolerance handling](#fault-tolerance) mechanisms, and\n- supports [distributed dataframes and distributed data loading](#distributed-data-loading)\n\nAll releases are tested on large clusters and workloads.\n\nThis package is based on \u003c!--$UNCOMMENT{ref}`XGBoost-Ray \u003cxgboost-ray\u003e`--\u003e\u003c!--$REMOVE--\u003e[XGBoost-Ray](https://github.com/ray-project/xgboost_ray)\u003c!--$END_REMOVE--\u003e. As of now, XGBoost-Ray is a dependency for LightGBM-Ray.\n\n## Installation\n\nYou can install the latest LightGBM-Ray release from PIP:\n\n```bash\npip install \"lightgbm_ray\"\n```\n\nIf you'd like to install the latest master, use this command instead:\n\n```bash\npip install \"git+https://github.com/ray-project/lightgbm_ray.git#egg=lightgbm_ray\"\n```\n\n## Usage\n\nLightGBM-Ray provides a drop-in replacement for LightGBM's `train`\nfunction. To pass data, a `RayDMatrix` object is required, common\nwith XGBoost-Ray. You can also use a scikit-learn\ninterface - see next section.\n\nJust as in original `lgbm.train()` function, the \n[training parameters](https://lightgbm.readthedocs.io/en/latest/Parameters.html)\nare passed as the `params` dictionary.\n\nRay-specific distributed training parameters are configured with a\n`lightgbm_ray.RayParams` object. For instance, you can set\nthe `num_actors` property to specify how many distributed actors\nyou would like to use.\n\nHere is a simplified example (which requires `sklearn`):\n\n**Training:**\n\n```python\nfrom lightgbm_ray import RayDMatrix, RayParams, train\nfrom sklearn.datasets import load_breast_cancer\n\ntrain_x, train_y = load_breast_cancer(return_X_y=True)\ntrain_set = RayDMatrix(train_x, train_y)\n\nevals_result = {}\nbst = train(\n    {\n        \"objective\": \"binary\",\n        \"metric\": [\"binary_logloss\", \"binary_error\"],\n    },\n    train_set,\n    evals_result=evals_result,\n    valid_sets=[train_set],\n    valid_names=[\"train\"],\n    verbose_eval=False,\n    ray_params=RayParams(num_actors=2, cpus_per_actor=2))\n\nbst.booster_.save_model(\"model.lgbm\")\nprint(\"Final training error: {:.4f}\".format(\n    evals_result[\"train\"][\"binary_error\"][-1]))\n```\n\n**Prediction:**\n\n```python\nfrom lightgbm_ray import RayDMatrix, RayParams, predict\nfrom sklearn.datasets import load_breast_cancer\nimport lightgbm as lgbm\n\ndata, labels = load_breast_cancer(return_X_y=True)\n\ndpred = RayDMatrix(data, labels)\n\nbst = lgbm.Booster(model_file=\"model.lgbm\")\npred_ray = predict(bst, dpred, ray_params=RayParams(num_actors=2))\n\nprint(pred_ray)\n```\n\n### scikit-learn API\n\nLightGBM-Ray also features a scikit-learn API fully mirroring pure\nLightGBM scikit-learn API, providing a completely drop-in\nreplacement. The following estimators are available:\n\n- `RayLGBMClassifier`\n- `RayLGBMRegressor`\n\nExample usage of `RayLGBMClassifier`:\n\n```python\nfrom lightgbm_ray import RayLGBMClassifier, RayParams\nfrom sklearn.datasets import load_breast_cancer\nfrom sklearn.model_selection import train_test_split\n\nseed = 42\n\nX, y = load_breast_cancer(return_X_y=True)\nX_train, X_test, y_train, y_test = train_test_split(\n    X, y, train_size=0.25, random_state=42)\n\nclf = RayLGBMClassifier(\n    n_jobs=2,  # In LightGBM-Ray, n_jobs sets the number of actors\n    random_state=seed)\n\n# scikit-learn API will automatically convert the data\n# to RayDMatrix format as needed.\n# You can also pass X as a RayDMatrix, in which case\n# y will be ignored.\n\nclf.fit(X_train, y_train)\n\npred_ray = clf.predict(X_test)\nprint(pred_ray)\n\npred_proba_ray = clf.predict_proba(X_test)\nprint(pred_proba_ray)\n\n# It is also possible to pass a RayParams object\n# to fit/predict/predict_proba methods - will override\n# n_jobs set during initialization\n\nclf.fit(X_train, y_train, ray_params=RayParams(num_actors=2))\n\npred_ray = clf.predict(X_test, ray_params=RayParams(num_actors=2))\nprint(pred_ray)\n```\n\nThings to keep in mind:\n\n- `n_jobs` parameter controls the number of actors spawned.\nYou can pass a `RayParams` object to the\n`fit`/`predict`/`predict_proba` methods as the `ray_params` argument \nfor greater control over resource allocation. Doing\nso will override the value of `n_jobs` with the value of\n`ray_params.num_actors` attribute. For more information, refer\nto the [Resources](#resources) section below.\n- By default `n_jobs` is set to `1`, which means the training\nwill **not** be distributed. Make sure to either set `n_jobs`\nto a higher value or pass a `RayParams` object as outlined above\nin order to take advantage of LightGBM-Ray's functionality.\n- After calling `fit`, additional evaluation results (e.g. training time,\nnumber of rows, callback results) will be available under\n`additional_results_` attribute.\n- `eval_` arguments are supported, but early stopping is not.\n- LightGBM-Ray's scikit-learn API is based on LightGBM 3.2.1.\nWhile we try to support older LightGBM versions, please note that\nthis library is only fully tested and supported for LightGBM \u003e= 3.2.1.\n\nFor more information on the scikit-learn API, refer to the [LightGBM documentation](https://lightgbm.readthedocs.io/en/latest/Python-API.html#scikit-learn-api).\n\n## Data loading\n\nData is passed to LightGBM-Ray via a `RayDMatrix` object.\n\nThe `RayDMatrix` lazy loads data and stores it sharded in the\nRay object store. The Ray LightGBM actors then access these\nshards to run their training on. \n\nA `RayDMatrix` support various data and file types, like\nPandas DataFrames, Numpy Arrays, CSV files and Parquet files.\n\nExample loading multiple parquet files:\n\n```python\nimport glob\nfrom lightgbm_ray import RayDMatrix, RayFileType\n\n# We can also pass a list of files\npath = list(sorted(glob.glob(\"/data/nyc-taxi/*/*/*.parquet\")))\n\n# This argument will be passed to `pd.read_parquet()`\ncolumns = [\n    \"passenger_count\",\n    \"trip_distance\", \"pickup_longitude\", \"pickup_latitude\",\n    \"dropoff_longitude\", \"dropoff_latitude\",\n    \"fare_amount\", \"extra\", \"mta_tax\", \"tip_amount\",\n    \"tolls_amount\", \"total_amount\"\n]\n\ndtrain = RayDMatrix(\n    path, \n    label=\"passenger_count\",  # Will select this column as the label\n    columns=columns,\n    # ignore=[\"total_amount\"],  # Optional list of columns to ignore\n    filetype=RayFileType.PARQUET)\n```\n\n\u003c!--$UNCOMMENT(lightgbm-ray-tuning)=--\u003e\n\n## Hyperparameter Tuning\n\nLightGBM-Ray integrates with  \u003c!--$UNCOMMENT{ref}`Ray Tune \u003ctune-main\u003e`--\u003e\u003c!--$REMOVE--\u003e[Ray Tune](https://tune.io)\u003c!--$END_REMOVE--\u003e to provide distributed hyperparameter tuning for your\ndistributed LightGBM models. You can run multiple LightGBM-Ray training runs in parallel, each with a different\nhyperparameter configuration, and each training run parallelized by itself. All you have to do is move your training\ncode to a function, and pass the function to `tune.run`. Internally, `train` will detect if `tune` is being used and will\nautomatically report results to tune.\n\nExample using LightGBM-Ray with Ray Tune:\n\n```python\nfrom lightgbm_ray import RayDMatrix, RayParams, train\nfrom sklearn.datasets import load_breast_cancer\n\nnum_actors = 2\nnum_cpus_per_actor = 2\n\nray_params = RayParams(\n    num_actors=num_actors, cpus_per_actor=num_cpus_per_actor)\n\ndef train_model(config):\n    train_x, train_y = load_breast_cancer(return_X_y=True)\n    train_set = RayDMatrix(train_x, train_y)\n\n    evals_result = {}\n    bst = train(\n        params=config,\n        dtrain=train_set,\n        evals_result=evals_result,\n        valid_sets=[train_set],\n        valid_names=[\"train\"],\n        verbose_eval=False,\n        ray_params=ray_params)\n    bst.booster_.save_model(\"model.lgbm\")\n\nfrom ray import tune\n\n# Specify the hyperparameter search space.\nconfig = {\n    \"objective\": \"binary\",\n    \"metric\": [\"binary_logloss\", \"binary_error\"],\n    \"eta\": tune.loguniform(1e-4, 1e-1),\n    \"subsample\": tune.uniform(0.5, 1.0),\n    \"max_depth\": tune.randint(1, 9)\n}\n\n# Make sure to use the `get_tune_resources` method to set the `resources_per_trial`\nanalysis = tune.run(\n    train_model,\n    config=config,\n    metric=\"train-binary_error\",\n    mode=\"min\",\n    num_samples=4,\n    resources_per_trial=ray_params.get_tune_resources())\nprint(\"Best hyperparameters\", analysis.best_config)\n```\n\nAlso see examples/simple_tune.py for another example.\n\n## Fault tolerance\n\nLightGBM-Ray leverages the stateful Ray actor model to\nenable fault tolerant training. Currently, only non-elastic\ntraining is supported.\n\n### Non-elastic training (warm restart)\n\nWhen an actor or node dies, LightGBM-Ray will retain the\nstate of the remaining actors. In non-elastic training,\nthe failed actors will be replaced as soon as resources\nare available again. Only these actors will reload their\nparts of the data. Training will resume once all actors\nare ready for training again.\n\nYou can configure this mode in the `RayParams`:\n\n```python\nfrom lightgbm_ray import RayParams\n\nray_params = RayParams(\n    max_actor_restarts=2,    # How often are actors allowed to fail, Default = 0\n)\n```\n\n## Resources\n\nBy default, LightGBM-Ray tries to determine the number of CPUs\navailable and distributes them evenly across actors.\n\nIn the case of very large clusters or clusters with many different\nmachine sizes, it makes sense to limit the number of CPUs per actor\nby setting the `cpus_per_actor` argument. Consider always\nsetting this explicitly.\n\nThe number of LightGBM actors always has to be set manually with\nthe `num_actors` argument.\n\n### Multi GPU training\n\nBy default, LightGBM-Ray tries to determine the number of CPUs\navailable and distributes them evenly across actors.\n\nIt is important to note that distributed LightGBM needs at least\ntwo CPUs per actor to function efficiently (without blocking).\nTherefore, by default, at least two CPUs will be assigned to each actor,\nand an exception will be raised if an actor has less than two CPUs.\nIt is possible to override this check by setting the\n`allow_less_than_two_cpus` argument to `True`, though it is not\nrecommended, as it will negatively impact training performance.\n\nIn the case of very large clusters or clusters with many different\nmachine sizes, it makes sense to limit the number of CPUs per actor\nby setting the `cpus_per_actor` argument. Consider always\nsetting this explicitly.\n\nThe number of LightGBM actors always has to be set manually with\nthe `num_actors` argument.\n\n### Multi GPU training\nLightGBM-Ray enables multi GPU training. The LightGBM core backend\nwill automatically handle communication.\nAll you have to do is to start one actor per GPU and set LightGBM's\n`device_type` to a GPU-compatible option, eg. `gpu` (see LightGBM\ndocumentation for more details.) \n\nFor instance, if you have 2 machines with 4 GPUs each, you will want\nto start 8 remote actors, and set `gpus_per_actor=1`. There is usually\nno benefit in allocating less (e.g. 0.5) or more than one GPU per actor. \n\nYou should divide the CPUs evenly across actors per machine, so if your \nmachines have 16 CPUs in addition to the 4 GPUs, each actor should have\n4 CPUs to use.\n\n```python\nfrom lightgbm_ray import RayParams\n\nray_params = RayParams(\n    num_actors=8,\n    gpus_per_actor=1,\n    cpus_per_actor=4,   # Divide evenly across actors per machine\n)\n```\n\n### How many remote actors should I use?\n\nThis depends on your workload and your cluster setup.\nGenerally there is no inherent benefit of running more than\none remote actor per node for CPU-only training. This is because\nLightGBM core can already leverage multiple CPUs via threading.\n\nHowever, there are some cases when you should consider starting\nmore than one actor per node:\n\n- For [**multi GPU training**](#multi-gpu-training), each GPU should have a separate\n  remote actor. Thus, if your machine has 24 CPUs and 4 GPUs,\n  you will want to start 4 remote actors with 6 CPUs and 1 GPU\n  each\n- In a **heterogeneous cluster**, you might want to find the\n  [greatest common divisor](https://en.wikipedia.org/wiki/Greatest_common_divisor)\n  for the number of CPUs.\n  E.g. for a cluster with three nodes of 4, 8, and 12 CPUs, respectively,\n  you should set the number of actors to 6 and the CPUs per \n  actor to 4.\n\n## Distributed data loading\n\nLightGBM-Ray can leverage both centralized and distributed data loading.\n\nIn **centralized data loading**, the data is partitioned by the head node\nand stored in the object store. Each remote actor then retrieves their\npartitions by querying the Ray object store. Centralized loading is used\nwhen you pass centralized in-memory dataframes, such as Pandas dataframes\nor Numpy arrays, or when you pass a single source file, such as a single CSV\nor Parquet file.\n\n\n```python\nfrom lightgbm_ray import RayDMatrix\n\n# This will use centralized data loading, as only one source file is specified\n# `label_col` is a column in the CSV, used as the target label\nray_params = RayDMatrix(\"./source_file.csv\", label=\"label_col\")\n```\n\nIn **distributed data loading**, each remote actor loads their data directly from\nthe source (e.g. local hard disk, NFS, HDFS, S3), \nwithout a central bottleneck. The data is still stored in the\nobject store, but locally to each actor. This mode is used automatically\nwhen loading data from multiple CSV or Parquet files. Please note that\nwe do not check or enforce partition sizes in this case - it is your job\nto make sure the data is evenly distributed across the source files.\n\n```python\nfrom lightgbm_ray import RayDMatrix\n\n# This will use distributed data loading, as four source files are specified\n# Please note that you cannot schedule more than four actors in this case.\n# `label_col` is a column in the Parquet files, used as the target label\nray_params = RayDMatrix([\n    \"hdfs:///tmp/part1.parquet\",\n    \"hdfs:///tmp/part2.parquet\",\n    \"hdfs:///tmp/part3.parquet\",\n    \"hdfs:///tmp/part4.parquet\",\n], label=\"label_col\")\n```\n\nLastly, LightGBM-Ray supports **distributed dataframe** representations, such\nas \u003c!--$UNCOMMENT{ref}`Ray Datasets \u003cdatasets\u003e`--\u003e\u003c!--$REMOVE--\u003e[Ray Datasets](https://docs.ray.io/en/latest/data/dataset.html)\u003c!--$END_REMOVE--\u003e,\n[Modin](https://modin.readthedocs.io/en/latest/) and\n[Dask dataframes](https://docs.dask.org/en/latest/dataframe.html)\n(used with \u003c!--$UNCOMMENT{ref}`Dask on Ray \u003cdask-on-ray\u003e`--\u003e\u003c!--$REMOVE--\u003e[Dask on Ray](https://docs.ray.io/en/master/dask-on-ray.html)\u003c!--$END_REMOVE--\u003e).\nHere, LightGBM-Ray will check on which nodes the distributed partitions \nare currently located, and will assign partitions to actors in order to\nminimize cross-node data transfer. Please note that we also assume here\nthat partition sizes are uniform. \n\n```python\nfrom lightgbm_ray import RayDMatrix\n\n# This will try to allocate the existing Modin partitions\n# to co-located Ray actors. If this is not possible, data will\n# be transferred across nodes\nray_params = RayDMatrix(existing_modin_df)\n```\n\n### Data sources\n\nThe following data sources can be used with a `RayDMatrix` object.\n\n| Type                                                             | Centralized loading | Distributed loading |\n|------------------------------------------------------------------|---------------------|---------------------|\n| Numpy array                                                      | Yes                 | No                  |\n| Pandas dataframe                                                 | Yes                 | No                  |\n| Single CSV                                                       | Yes                 | No                  |\n| Multi CSV                                                        | Yes                 | Yes                 |\n| Single Parquet                                                   | Yes                 | No                  |\n| Multi Parquet                                                    | Yes                 | Yes                 |\n| [Ray Dataset](https://docs.ray.io/en/latest/data/dataset.html)   | Yes                 | Yes                 |\n| [Petastorm](https://github.com/uber/petastorm)                   | Yes                 | Yes                 |\n| [Dask dataframe](https://docs.dask.org/en/latest/dataframe.html) | Yes                 | Yes                 |\n| [Modin dataframe](https://modin.readthedocs.io/en/latest/)       | Yes                 | Yes                 |\n\n## Memory usage\n\nDetails coming soon.\n\u003c!-- This hasn't been verifiec --\u003e\n\u003c!-- \nXGBoost uses a compute-optimized datastructure, the `DMatrix`,\nto hold training data. When converting a dataset to a `DMatrix`,\nXGBoost creates intermediate copies and ends up \nholding a complete copy of the full data. The data will be converted\ninto the local dataformat (on a 64 bit system these are 64 bit floats.)\nDepending on the system and original dataset dtype, this matrix can \nthus occupy more memory than the original dataset.\n\nThe **peak memory usage** for CPU-based training is at least\n**3x** the dataset size (assuming dtype `float32` on a 64bit system) \nplus about **400,000 KiB** for other resources,\nlike operating system requirements and storing of intermediate\nresults.\n\n**Example**\n- Machine type: AWS m5.xlarge (4 vCPUs, 16 GiB RAM)\n- Usable RAM: ~15,350,000 KiB\n- Dataset: 1,250,000 rows with 1024 features, dtype float32.\n  Total size: 5,000,000 KiB\n- XGBoost DMatrix size: ~10,000,000 KiB\n\nThis dataset will fit exactly on this node for training.\n\nNote that the DMatrix size might be lower on a 32 bit system. \n\n**GPUs**\n\nGenerally, the same memory requirements exist for GPU-based\ntraining. Additionally, the GPU must have enough memory\nto hold the dataset. \n\nIn the example above, the GPU must have at least \n10,000,000 KiB (about 9.6 GiB) memory. However, \nempirically we found that using a `DeviceQuantileDMatrix`\nseems to show more peak GPU memory usage, possibly \nfor intermediate storage when loading data (about 10%). --\u003e\n\n**Best practices**\n\nIn order to reduce peak memory usage, consider the following\nsuggestions:\n\n- Store data as `float32` or less. More precision is often \n  not needed, and keeping data in a smaller format will\n  help reduce peak memory usage for initial data loading.\n- Pass the `dtype` when loading data from CSV. Otherwise,\n  floating point values will be loaded as `np.float64` \n  per default, increasing peak memory usage by 33%.\n\n## Placement Strategies\n\nLightGBM-Ray leverages Ray's Placement Group API (https://docs.ray.io/en/master/placement-group.html)\nto implement placement strategies for better fault tolerance. \n\nBy default, a SPREAD strategy is used for training, which attempts to spread all of the training workers\nacross the nodes in a cluster on a best-effort basis. This improves fault tolerance since it minimizes the \nnumber of worker failures when a node goes down, but comes at a cost of increased inter-node communication\nTo disable this strategy, set the `RXGB_USE_SPREAD_STRATEGY` environment variable to 0. If disabled, no\nparticular placement strategy will be used.\n\n\u003c!-- Note that this strategy is used only when `elastic_training` is not used. If `elastic_training` is set to `True`,\nno placement strategy is used. --\u003e\n\nWhen LightGBM-Ray is used with Ray Tune for hyperparameter tuning, a PACK strategy is used. This strategy\nattempts to place all workers for each trial on the same node on a best-effort basis. This means that if a node\ngoes down, it will be less likely to impact multiple trials.\n\nWhen placement strategies are used, LightGBM-Ray will wait for 100 seconds for the required resources\nto become available, and will fail if the required resources cannot be reserved and the cluster cannot autoscale\nto increase the number of resources. You can change the `RXGB_PLACEMENT_GROUP_TIMEOUT_S` environment variable to modify \nhow long this timeout should be. \n\n## More examples\n\nFor complete end to end examples, please have a look at\nthe [examples folder](https://github.com/ray-project/lightgbm_ray/tree/main/lightgbm_ray/examples/):\n\n* [Simple sklearn breastcancer dataset example](https://github.com/ray-project/lightgbm_ray/tree/main/lightgbm_ray/examples/simple.py) (requires `sklearn`)\n* [HIGGS classification example](https://github.com/ray-project/lightgbm_ray/tree/main/lightgbm_ray/examples/higgs.py) \n([download dataset (2.6 GB)](https://archive.ics.uci.edu/ml/machine-learning-databases/00280/HIGGS.csv.gz))\n* [HIGGS classification example with Parquet](https://github.com/ray-project/lightgbm_ray/tree/main/lightgbm_ray/examples/higgs_parquet.py) (uses the same dataset) \n* [Test data classification](https://github.com/ray-project/lightgbm_ray/tree/main/lightgbm_ray/examples/train_on_test_data.py) (uses a self-generated dataset)\n\u003c!--$REMOVE--\u003e\n## Resources\n\n* [LightGBM-Ray documentation](https://docs.ray.io/en/master/lightgbm-ray.html)\n* [Ray community slack](https://forms.gle/9TSdDYUgxYs8SA9e8)\n\u003c!--$END_REMOVE--\u003e\n\u003c!--$UNCOMMENT## API reference\n\n```{eval-rst}\n.. autoclass:: lightgbm_ray.RayParams\n    :members:\n```\n\n```{eval-rst}\n.. note::\n  The ``xgboost_ray.RayDMatrix`` class is shared with :ref:`XGBoost-Ray \u003cxgboost-ray\u003e`.\n\n.. autoclass:: xgboost_ray.RayDMatrix\n    :members:\n    :noindex:\n```\n\n```{eval-rst}\n.. autofunction:: lightgbm_ray.train\n```\n\n```{eval-rst}\n.. autofunction:: lightgbm_ray.predict\n```\n\n### scikit-learn API\n\n```{eval-rst}\n.. autoclass:: lightgbm_ray.RayLGBMClassifier\n    :members:\n```\n\n```{eval-rst}\n.. autoclass:: lightgbm_ray.RayLGBMRegressor\n    :members:\n```--\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fray-project%2Flightgbm_ray","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fray-project%2Flightgbm_ray","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fray-project%2Flightgbm_ray/lists"}