{"id":30598257,"url":"https://github.com/redhat-na-ssa/nvidia-vgpu-openshift","last_synced_at":"2026-02-10T09:32:29.864Z","repository":{"id":294482318,"uuid":"987116693","full_name":"redhat-na-ssa/nvidia-vgpu-openshift","owner":"redhat-na-ssa","description":"Making vGPUs work with OpenShift Virtualization","archived":false,"fork":false,"pushed_at":"2025-07-18T19:38:24.000Z","size":227,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":6,"default_branch":"main","last_synced_at":"2025-07-19T00:14:12.145Z","etag":null,"topics":["gpus","nvidia","openshift","vgpu"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/redhat-na-ssa.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-05-20T15:44:02.000Z","updated_at":"2025-07-18T19:38:28.000Z","dependencies_parsed_at":"2025-07-18T21:19:54.171Z","dependency_job_id":"5f39d4c6-b2a4-47bf-92be-fd202160f85b","html_url":"https://github.com/redhat-na-ssa/nvidia-vgpu-openshift","commit_stats":null,"previous_names":["redhat-na-ssa/nvidia-vgpu-openshift"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/redhat-na-ssa/nvidia-vgpu-openshift","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/redhat-na-ssa%2Fnvidia-vgpu-openshift","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/redhat-na-ssa%2Fnvidia-vgpu-openshift/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/redhat-na-ssa%2Fnvidia-vgpu-openshift/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/redhat-na-ssa%2Fnvidia-vgpu-openshift/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/redhat-na-ssa","download_url":"https://codeload.github.com/redhat-na-ssa/nvidia-vgpu-openshift/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/redhat-na-ssa%2Fnvidia-vgpu-openshift/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":272772757,"owners_count":24990514,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-08-29T02:00:10.610Z","response_time":87,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["gpus","nvidia","openshift","vgpu"],"created_at":"2025-08-29T22:13:04.657Z","updated_at":"2026-02-10T09:32:29.802Z","avatar_url":"https://github.com/redhat-na-ssa.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# NVIDIA vGPU on OpenShift\n\nThis repository serves as a complement to the [NVIDIA documentation on vGPUs on OpenShift](https://docs.nvidia.com/datacenter/cloud-native/openshift/latest/openshift-virtualization.html) to make it easier to deploy the solution on a OpenShift cluster\n\nThere are two ways to configure mediated devices in OpenShift, one is with the OpenShift Virtualization operator and the other is with the NVIDIA GPU operator, this doc focuses on the latter. [Docs steps for configuring with OpenShift Virtualization](https://docs.redhat.com/en/documentation/openshift_container_platform/4.18/html/virtualization/managing-vms#virt-options-configuring-mdevs_virt-configuring-virtual-gpus)\n\n\n## Testing environment\n\nThis repository has been tested on:\n\n* OpenShift 4.18.13 cluster (stable-4.18 channel)\n  * NVIDIA GPU Operator v25.3.0\n  * NVIDIA Driver 570.133.10\n  * NVIDIA T4 GPU (AWS instance [g4dn.metal](https://instances.vantage.sh/aws/ec2/g4dn.metal))\n\n## Prerequisites\n\nThe NVIDIA vGPU software needs to be obtained from the [NVIDIA Licensing Portal](https://nvid.nvidia.com/dashboard/#/dashboard):\n\n(Following block is replicated from the NVIDIA documentation)\n\n\u003e [!NOTE]\n\u003e Login to the NVIDIA Licensing Portal and navigate to the Software Downloads section.\n\u003e\n\u003e The NVIDIA vGPU Software is located on the Driver downloads tab of the Software Downloads page.\n\u003e\n\u003e Click the Download link for the Linux KVM complete vGPU package. Confirm that the Product Version column shows the vGPU version to \u003e install. Unzip the bundle to obtain the NVIDIA vGPU Manager for Linux file, NVIDIA-Linux-x86_64-\u003cversion\u003e-vgpu-kvm.run.\n\u003e\n\u003e NVIDIA AI Enterprise customers must use the aie .run file for building the NVIDIA vGPU Manager image. Download the NVIDIA-Linux-x86_64-\u003cversion\u003e-vgpu-kvm-aie.run file instead, and rename it to NVIDIA-Linux-x86_64-\u003cversion\u003e-vgpu-kvm.run before proceeding with the rest of the procedure.\n\n## Instructions\n\n### Build Driver image on local machine\n\nDisconnected vgpu-manager [build documentation](DISCONNECTED-vgpu-manager.md)\n\nSince we need the proprietary drivers from NVIDIA, the easiest way to get started is by building the container image locally\n\nLocal requirements:\n\n* git\n* podman\n\n```bash\ngit clone https://gitlab.com/nvidia/container-images/driver\ncd driver/vgpu-manager/rhel8/\ncp path/to/NVIDIA-Linux-x86_64-\u003cversion\u003e-vgpu-kvm.run .\n\n# the DRIVER_VERSION buildarg is from driver filename: NVIDIA-Linux-x86_64-570.133.10-vgpu-kvm.run\npodman build --build-arg DRIVER_VERSION=570.133.10 -t nvidia-driver:latest .\n```\n\n### (Optional) Push to internal registry\n\nOnce we have the container image we can use the OpenShift image registry to host but you can also use another registry if available\n\nIn the example we are using the driver version and the version of OpenShift in the tag in the form \"driver_version-rhcos_version\": `570.133.10-rhcos4.18`\n\n```bash\noc login https://api.cluster.example.com:6443 --username admin\noc registry login\n\npodman tag nvidia-driver:latest image-registry.openshift-image-registry.svc:5000/nvidia-gpu-operator/vgpu-manager:570.133.10-rhcos4.18\npodman push --tls-verify=false image-registry.openshift-image-registry.svc:5000/nvidia-gpu-operator/vgpu-manager:570.133.10-rhcos4.18\n```\n\n## Deploy GPU Operator\n\n### Label the nodes\n\nNodes that will be used for vGPU should be labeled to ensure that the GPU operator installs the correct NVIDIA drivers\n\n\u003e [!WARNING]\nIf the workload label was unset or set incorrectly you will have to remove the kernel modules that are loaded in order to change the supported workload for the node (rebooting the node is sufficient).\n\n(from the NVIDIA documentation)\n\u003e If the node label nvidia.com/gpu.workload.config does not exist on the node, the GPU Operator assumes the default GPU workload configuration, container, and deploys the software components needed to support this workload type. To change the default GPU workload configuration, set the following value in ClusterPolicy: .sandboxWorkloads.defaultWorkload=\u003cconfig\u003e.\n\nThe options are:\n* container\n* vm-passthrough\n* vm-vgpu\n\n```sh\noc label node ip-10-0-2-117.us-east-2.compute.internal --overwrite nvidia.com/gpu.workload.config=vm-vgpu\n```\n\n```sh\nnode/ip-10-0-2-117.us-east-2.compute.internal labeled\n```\n\nCheck the full list of nodes for the labels, in this case only the first node will be set up for vGPU, the rest will get the standard container-based workload support for GPUs\n\n```sh\noc get nodes -o custom-columns=Name:.metadata.name,GPU:.metadata.labels.'nvidia\\.com/gpu\\.workload\\.config'\n```\n\n```sh\nName                                        GPU\nip-10-0-2-117.us-east-2.compute.internal    vm-vgpu\nip-10-0-42-92.us-east-2.compute.internal    \u003cnone\u003e\nip-10-0-5-156.us-east-2.compute.internal    \u003cnone\u003e\n```\n\n### Install the vGPU manager image\n\n#### With CLI\n\n\u003e [!NOTE]\n\u003e If you are using an external registry, change the registry URI in the YAML before applying\n\nApply the ClusterPolicy in this repository\n\n```sh\noc apply -f clusterpolicy.yaml\n```\n\n#### With GUI\n\nOnce installing the GPU operator, add a new ClusterPolicy and switch to YAML\n\n![clusterpolicy](images/clusterpolicy-create.png)\n![clusterpolicy](images/clusterpolicy-yaml.png)\n\n\u003e [!NOTE]\n\u003e Modify the YAML and set the repository to the correct URL if you are not using the internal registry\n\nFind the settings listed below and make sure they are set, leave everything else as the default\n\n```sh\napiVersion: nvidia.com/v1\nkind: ClusterPolicy\nmetadata:\n  name: gpu-cluster-policy\nspec:\n  sandboxWorkloads:\n    enabled: true\n  vgpuManager:\n    enabled: true\n    image: vgpu-manager\n    repository: image-registry.openshift-image-registry.svc:5000/nvidia-gpu-operator\n    version: 570.133.10\n...\n```\n\n### Verification\n\n#### Check the node is advertising GPUs\n\n```sh\noc get nodes ip-10-0-2-117.us-east-2.compute.internal -o json | jq .status.allocatable\n```\n\n```json\n{\n \"cpu\": \"95500m\",\n \"devices.kubevirt.io/kvm\": \"1k\",\n \"devices.kubevirt.io/tun\": \"1k\",\n \"devices.kubevirt.io/vhost-net\": \"1k\",\n \"ephemeral-storage\": \"191655242229\",\n \"hugepages-1Gi\": \"0\",\n \"hugepages-2Mi\": \"0\",\n \"memory\": \"394701388Ki\",\n \"nvidia.com/GRID_T4-8Q\": \"16\",\n \"nvidia.com/gpu\": \"0\",\n \"pods\": \"250\"\n}\n```\n\n## Configure OpenShift virtualization\n\n### Modify Hyperconverged object\n\nThe `resourceName` should correspond to the name of the resource from the allocatable resources, in this example `nvidia.com/GRID_T4-8Q`\n\n```yaml\napiVersion: hco.kubevirt.io/v1beta1\nkind: HyperConverged\nspec:\n  featureGates:\n    disableMDevConfiguration: true\n  permittedHostDevices:\n    mediatedDevices:\n    - externalResourceProvider: true\n      mdevNameSelector: GRID T4-8Q\n      resourceName: nvidia.com/GRID_T4-8Q\n```\n\nAfter which, you should be able to create or modify a VM and add a GPU device.\n\n## Use a custom vGPU configuration\n\nSo far we have been using the default configuration. If you want to set all GPUs in the node to a specific configuration you can just use the `nvidia.com/vgpu.config` label. However to do a custom configuration per GPU in a cluster you need to apply a ConfigMap and then reference it in the ClusterPolicy.\n\nExample in this repo:\n\n```\noc apply -f configmap.yaml\n```\n\n```\noc patch clusterpolicy gpu-cluster-policy --type=merge -p '{\"spec\":{\"vgpuDeviceManager\":{\"config\":{\"name\":\"custom-vgpu-config\"}}}}'\n```\n\nThen you can apply the configuration to a node:\n\n```\noc label node ip-10-0-31-236.us-east-2.compute.internal nvidia.com/vgpu.config=T4-custom\n```\n\n```\noc get nodes ip-10-0-31-236.us-east-2.compute.internal -o yaml | grep -A 12 allocatable| grep nvidia\n```\n```\n    nvidia.com/GRID_T4-8C: \"2\"\n    nvidia.com/GRID_T4-8Q: \"12\"\n    nvidia.com/GRID_T4-16C: \"1\"\n```\n\n## Testing in the VM\n\nThe VM needs to use the GRID drivers which are used to communicate with the vGPU host driver. For this test I used `nvidia-linux-grid-570-570.133.20-1.x86_64.rpm`\n\nEPEL is required for DKMS:\n\n```\ndnf install https://dl.fedoraproject.org/pub/epel/epel-release-latest-9.noarch.rpm\ndnf install nvidia-linux-grid-570-570.133.20-1.x86_64.rpm\n```\n\nAfter which, we have the correct GPU available:\n\n```\nnvidia-smi\n```\n```\nWed May 28 09:44:12 2025\n+-----------------------------------------------------------------------------------------+\n| NVIDIA-SMI 570.133.20             Driver Version: 570.133.20     CUDA Version: 12.8     |\n|-----------------------------------------+------------------------+----------------------+\n| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |\n| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |\n|                                         |                        |               MIG M. |\n|=========================================+========================+======================|\n|   0  GRID T4-16C                    On  |   00000000:09:00.0 Off |                    0 |\n| N/A   N/A    P8            N/A  /  N/A  |       1MiB /  16384MiB |      0%      Default |\n|                                         |                        |                  N/A |\n+-----------------------------------------+------------------------+----------------------+\n\n+-----------------------------------------------------------------------------------------+\n| Processes:                                                                              |\n|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |\n|        ID   ID                                                               Usage      |\n|=========================================================================================|\n|  No running processes found                                                             |\n+-----------------------------------------------------------------------------------------+\n```\n\n### Testing with PyTorch:\n\nAdding cuda-toolkit per docs: https://developer.nvidia.com/cuda-downloads\n\n```\ndnf config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel9/x86_64/cuda-rhel9.repo\ndnf -y install cuda-toolkit-12-9\n```\n\nCreating a virtual environment and installing pytorch:\n\n```\npython3 -m venv env\nenv/bin/pip install torch\nenv/bin/pip install numpy\n```\n\n`test_torch.py`\n```\nimport sys\nimport torch\n\ndef main():\n    print(\"Python Version:\", sys.version)\n    print(\"Torch Version:\", torch.__version__)\n    print(\"CUDA Available:\", torch.cuda.is_available())\n    print(\"CUDA Device Count:\", torch.cuda.device_count())\n\n    if torch.cuda.is_available():\n        print(\"CUDA Device Name:\", torch.cuda.get_device_name(0))\n\nif __name__ == '__main__':\n    main()\n```\n\n```\nenv/bin/python test_torch.py\n```\n```\nPython Version: 3.9.21 (main, Feb 10 2025, 00:00:00)\n[GCC 11.5.0 20240719 (Red Hat 11.5.0-5)]\nTorch Version: 2.7.0+cu126\nCUDA Available: True\nCUDA Device Count: 1\nCUDA Device Name: GRID T4-16C\n```\n\n## Other Verification and Debugging steps\n\n### gpu-operator must-gather\n\n```\ncurl -o must-gather.sh -L https://raw.githubusercontent.com/NVIDIA/gpu-operator/main/hack/must-gather.sh\nchmod +x must-gather.sh\n./must-gather.sh\n```\n\n### Node labels\n\nInitially the gpu operator should label the node with:\n\n```\n\"nvidia.com/gpu.deploy.cc-manager\": \"true\",\n\"nvidia.com/gpu.deploy.nvsm\": \"\",\n\"nvidia.com/gpu.deploy.sandbox-device-plugin\": \"true\",\n\"nvidia.com/gpu.deploy.sandbox-validator\": \"true\",\n\"nvidia.com/gpu.deploy.vgpu-device-manager\": \"true\",\n\"nvidia.com/gpu.deploy.vgpu-manager\": \"true\",\n\"nvidia.com/gpu.present\": \"true\",\n```\n\nAfter everything installs it adds\n\n```\n\"nvidia.com/vgpu.config.state\": \"success\",\n```\n\n#### Delete the node labels to reset\n\nThis is a bash function that removes all labels with the prefix `nvidia.com` from the node\n\n```\nnode_cleanup() {\n  local NODE=\"$1\"\n  oc get node \"$NODE\" -o go-template='{{range $key, $value := .metadata.labels}}{{if and (ge (len $key) 11) (eq (slice $key 0 11) \"nvidia.com/\")}}{{$key}}{{\"\\n\"}}{{end}}{{end}}' | xargs -t -I {} oc label node \"$NODE\" {}-\n}\n\nnode_cleanup \u003cnode-name\u003e\n```\n\n### Pods\n\nThe gpu operator should control four Daemonsets for deploying vGPU manager onto the node:\n\n```\n$ oc get ds -n nvidia-gpu-operator\nNAME                                                  NODE SELECTOR\nnvidia-sandbox-device-plugin-daemonset                nvidia.com/gpu.deploy.sandbox-device-plugin=true\nnvidia-sandbox-validator                              nvidia.com/gpu.deploy.sandbox-validator=true\nnvidia-vgpu-device-manager                            nvidia.com/gpu.deploy.vgpu-device-manager=true\nnvidia-vgpu-manager-daemonset-416.94.202505051351-0   feature.node.kubernetes.io/system-os_release.OSTREE_VERSION=416.94.202505051351-0,nvidia.com/gpu.deploy.vgpu-manager=true\n```\n\n### Driver version\n\nEnsure that the NVIDIA GPU drivers are the correct version\n\n```sh\noc debug node/ip-10-0-2-117.us-east-2.compute.internal\n```\n\n```sh\nStarting pod/ip-10-0-2-117us-east-2computeinternal-debug-pk9mk ...\nTo use host binaries, run `chroot /host`\nPod IP: 10.0.2.117\nIf you don't see a command prompt, try pressing enter.\n```\n\n```sh\ncat /sys/module/nvidia/version\n```\n\n```output\n570.133.10\n```\n\n#### Check PCI cards are present\n\n```sh\noc debug node/ip-10-0-2-117.us-east-2.compute.internal\n```\n\n```output\nStarting pod/ip-10-0-2-117us-east-2computeinternal-debug-pk9mk ...\nTo use host binaries, run `chroot /host`\nPod IP: 10.0.2.117\nIf you don't see a command prompt, try pressing enter.\nsh-5.1#\n```\n\n```sh\nlspci -nnk -d 10de:\n```\n\n```output\n18:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\nlspci: Unable to load libkmod resources: error -2\n19:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\n35:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\n36:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\ne7:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\ne8:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\nf4:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\nf5:00.0 3D controller [0302]: NVIDIA Corporation TU104GL [Tesla T4] [10de:1eb8] (rev a1)\n        Subsystem: NVIDIA Corporation Device [10de:12a2]\n        Kernel driver in use: nvidia\n```\n\n#### Get the supported mediated devices types\n\n```\nfor pci in $(lspci | grep -i nvidia | awk '{print $1}'); do  find /sys/bus/pci/devices/0000:$pci/mdev_supported_types/ -name name -exec cat {} \\;; done 2\u003e /dev/null | sort -u\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fredhat-na-ssa%2Fnvidia-vgpu-openshift","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fredhat-na-ssa%2Fnvidia-vgpu-openshift","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fredhat-na-ssa%2Fnvidia-vgpu-openshift/lists"}