Add kubeflow/v1.10.0

This commit is contained in:
wbsong111
2025-06-24 12:03:10 +09:00
parent 0132496142
commit 6e8dd89e48
1531 changed files with 120984 additions and 262120 deletions
+11
View File
@@ -0,0 +1,11 @@
KUBERAY_RELEASE_VERSION ?= 1.3.2
KUBERAY_HELM_CHART_REPO ?= https://ray-project.github.io/kuberay-helm/
.PHONY: kuberay-operator/base
kuberay-operator/base:
mkdir -p kuberay-operator/base
cd kuberay-operator/base && helm template --include-crds kuberay-operator kuberay-operator --version ${KUBERAY_RELEASE_VERSION} --repo ${KUBERAY_HELM_CHART_REPO} > resources.yaml
.PHONY: test
test:
./test.sh ${KF_PROFILE}
+5
View File
@@ -0,0 +1,5 @@
approvers:
- juliusvonkohout
reviewers:
- juliusvonkohout
- kimwnasptd
+150
View File
@@ -0,0 +1,150 @@
> Credit: This manifest refers a lot to the engineering blog ["Building a Machine Learning Platform with Kubeflow and Ray on Google Kubernetes Engine"](https://cloud.google.com/blog/products/ai-machine-learning/build-a-ml-platform-with-kubeflow-and-ray-on-gke) from Google Cloud.
# Ray
[Ray](https://github.com/ray-project/ray) is a unified framework for scaling AI and Python applications. Ray consists of a core distributed runtime and a toolkit of libraries (Ray AIR) for simplifying ML compute.
<figure>
<img
src="assets/map-of-ray.png"
alt="Ray">
<figcaption>Stack of Ray libraries - unified toolkit for ML workloads. (ref: https://docs.ray.io/en/latest/ray-overview/index.html)</figcaption>
</figure>
# KubeRay
[KubeRay](https://github.com/ray-project/kuberay) is an open-source Kubernetes operator for Ray. It provides several CRDs to simplify managing Ray clusters on Kubernetes. We will integrate Kubeflow and KubeRay in this document.
# Requirements
* Dependencies
* `kustomize`: v5.4.3+ (Kubeflow manifest is sensitive to `kustomize` version.)
* `Kubernetes`: v1.32+
* Computing resources:
* 16GB RAM
* 8 CPUs
# Example
<figure>
<img
src="assets/architecture.svg"
alt="ray/kubeflow integration">
<figcaption>Note: (1) Kubeflow Central Dashboard will be renamed to workbench in the future. (2) Kubeflow Pipeline (KFP) is an important component of Kubeflow, but it is not included in this example.</figcaption>
</figure>
## Step 1: Install Kubeflow
* This example installs Kubeflow with the master branch
* Install all Kubeflow official components and all common services using [one command](https://github.com/kubeflow/manifests/tree/master#install-with-a-single-command).
* If you do not want to install all components, you can comment out **KNative**, **Katib**, **Tensorboards Controller**, **Tensorboard Web App**, **Training Operator**, and **KServe** from [example/kustomization.yaml](https://github.com/kubeflow/manifests/blob/master/example/kustomization.yaml).
## Step 2: Install KubeRay operator
We never ever break Kubernetes standards and do not use the "default" namespace, but a proper one, in our case "kubeflow" for the ray operator.
```sh
# Install a KubeRay operator and custom resource definitions.
kustomize build kuberay-operator/overlays/kubeflow | kubectl apply --server-side -f -
# Check KubeRay operator
kubectl get pod -l app.kubernetes.io/component=kuberay-operator -n kubeflow
# NAME READY STATUS RESTARTS AGE
# kuberay-operator-5b8cd69758-rkpvh 1/1 Running 0 6m23s
```
> If you are creating a new namespace other than the kubeflow-user-example-com please follow below step otherwise skip the step.
## Step 3: Create a namespace
```sh
# Create a namespace: example-"development"
kubectl create ns development
# Enable istio-injection for the namespace
kubectl label namespace development istio-injection=enabled
# After creating the namespace, You have to do below mentioned changes in raycluster_example.yaml file(Required changes are also mentioned as comments in yaml file itself)
# 01. Replace the namesapce of AuthorizationPolicy principal
principals:
- "cluster.local/ns/development/sa/default-editor"
# 02. Replace the namespace of node-ip-address of headGroupSpec and workerGroupSpec
node-ip-address: $(hostname -I | tr -d ' ' | sed 's/\./-/g').raycluster-istio-headless-svc.development.svc.cluster.local
```
## Step 4: Install RayCluster
```sh
# Create a RayCluster CR, and the KubeRay operator will reconcile a Ray cluster
# with 1 head Pod and 1 worker Pod.
# $MY_KUBEFLOW_USER_NAMESPACE is the namespace that has been created in the above step.
export MY_KUBEFLOW_USER_NAMESPACE=development
kubectl apply -f raycluster_example.yaml -n $MY_KUBEFLOW_USER_NAMESPACE
# Check RayCluster
kubectl get pod -l ray.io/cluster=kubeflow-raycluster -n $MY_KUBEFLOW_USER_NAMESPACE
# NAME READY STATUS RESTARTS AGE
# kubeflow-raycluster-head-p6dpk 1/1 Running 0 70s
# kubeflow-raycluster-worker-small-group-l7j6c 1/1 Running 0 70s
#Check Raycluster headless service
kubectl get svc -n $MY_KUBEFLOW_USER_NAMESPACE
```
* `raycluster_example.yaml` uses `rayproject/ray:2.23.0-py311-cpu` as its OCI image. Ray is very sensitive to the Python versions and Ray versions between the server (RayCluster) and client (JupyterLab) sides. This image uses:
* Python 3.11
* Ray 2.23.0
## Step 5: Forward the port of Istio's Ingress-Gateway
* Follow the [instructions](https://github.com/kubeflow/manifests/tree/master#port-forward) to forward the port of Istio's Ingress-Gateway and log in to Kubeflow Central Dashboard.
## Step 6: Create a JupyterLab via Kubeflow Central Dashboard
* Click "Notebooks" icon in the left panel.
* Click "New Notebook"
* Select `kubeflownotebookswg/jupyter-scipy:v1.9.2` as OCI image (or any other with the same python version)
* Click "Launch"
* Click "CONNECT" to connect into the JupyterLab instance.
## Step 7: Use Ray client in the JupyterLab to connect to the RayCluster
* As I mentioned in Step 3, Ray is very sensitive to the Python versions and Ray versions between the server (RayCluster) and client (JupyterLab) sides.
```sh
# Check Python version. The version's MAJOR and MINOR should match with RayCluster (i.e. Python 3.11.9)
python --version
# Python 3.11.9
pip install -U ray[default]==2.23.0
```
* Connect to RayCluster via Ray client.
```python
# Open a new .ipynb page.
import ray
# For other namespaces use ray://${RAYCLUSTER_HEAD_SVC}.${NAMESPACE}.svc.cluster.local:${RAY_CLIENT_PORT}
# But we use of course our per namespace ray cluster to have multi-tenancy and
# We never ever use "default" as namespace since this would violate Kubernetes standards
ray.init(address="ray://kubeflow-raycluster-head-svc:10001")
print(ray.cluster_resources())
# {'node:10.244.0.41': 1.0, 'memory': 3000000000.0, 'node:10.244.0.40': 1.0, 'object_store_memory': 805386239.0, 'CPU': 2.0}
# Try Ray task
@ray.remote
def f(x):
return x * x
futures = [f.remote(i) for i in range(4)]
print(ray.get(futures)) # [0, 1, 4, 9]
# Try Ray actor
@ray.remote
class Counter(object):
def __init__(self):
self.n = 0
def increment(self):
self.n += 1
def read(self):
return self.n
counters = [Counter.remote() for i in range(4)]
[c.increment.remote() for c in counters]
futures = [c.read.remote() for c in counters]
print(ray.get(futures)) # [1, 1, 1, 1]
```
# Upgrading
See [UPGRADE.md](UPGRADE.md) for more details.
@@ -0,0 +1,6 @@
# Upgrading
```sh
# Step 1: Update KUBERAY_RELEASE_VERSION in Makefile
# Step 2: Create new KubeRay operator manifest
make kuberay-operator/base
```
File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 217 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 70 KiB

@@ -0,0 +1,52 @@
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kubeflow-kuberay-admin
labels:
app: kuberay-operator
app.kubernetes.io/name: kuberay-operator
rbac.authorization.kubeflow.org/aggregate-to-kubeflow-admin: "true"
rules: []
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kubeflow-kuberay-edit
labels:
app: kuberay-operator
app.kubernetes.io/name: kuberay-operator
rbac.authorization.kubeflow.org/aggregate-to-kubeflow-edit: "true"
rbac.authorization.kubeflow.org/aggregate-to-kubeflow-admin: "true"
rules:
- apiGroups:
- ray.io
resources:
- "*"
verbs:
- get
- list
- watch
- create
- update
- patch
- delete
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kubeflow-kuberay-view
labels:
app: kuberay-operator
app.kubernetes.io/name: kuberay-operator
rbac.authorization.kubeflow.org/aggregate-to-kubeflow-view: "true"
rules:
- apiGroups:
- ray.io
resources:
- "*"
verbs:
- get
- list
- watch
---
@@ -0,0 +1,9 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: kubeflow
resources:
- resources.yaml
- aggregated-roles.yaml
# The securitycontext for PSS restricted is now directly provided in the upstream manifests
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,13 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: kuberay-operator
namespace: kubeflow
spec:
template:
spec:
containers:
- name: kuberay-operator
env:
- name: ENABLE_INIT_CONTAINER_INJECTION
value: "false" # TODO maybe we can drop that with istio native sidecars
@@ -0,0 +1,6 @@
namespace: kubeflow
resources:
- ../../base
patches:
- path: disable-injection.yaml
@@ -0,0 +1,3 @@
namespace: kubeflow
resources:
- ../../base
@@ -0,0 +1,167 @@
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
name: allow-ray-workers-head
spec:
action: ALLOW
rules:
- from:
- source:
principals:
# kubeflow-user-example-com should be replaced with the namespace where the Ray cluster is being deployed
# TODO automatically use the current namespace
- "cluster.local/ns/kubeflow-user-example-com/sa/default-editor"
- to:
- operation:
ports:
- "6379"
- "6380"
- "6381"
- "6382"
- "6383"
- "52365"
- "8080"
- "10012"
---
apiVersion: v1
kind: Service
metadata:
labels:
ray.io/headless-worker-svc: raycluster-istio
name: raycluster-istio-headless-svc
spec:
clusterIP: None
selector:
ray.io/cluster: kubeflow-raycluster
publishNotReadyAddresses: true
ports:
- name: node-manager-port
port: 6380
appProtocol: grpc
- name: object-manager-port
port: 6381
appProtocol: grpc
- name: runtime-env-agent-port
port: 6382
appProtocol: grpc
- name: dashboard-agent-grpc-port
port: 6383
appProtocol: grpc
- name: dashboard-agent-listen-port
port: 52365
appProtocol: http
- name: metrics-export-port
port: 8080
appProtocol: http
- name: max-worker-port
port: 10012
appProtocol: grpc
---
apiVersion: ray.io/v1
kind: RayCluster
metadata:
name: kubeflow-raycluster
spec:
rayVersion: '2.44.1'
enableInTreeAutoscaling: true
autoscalerOptions:
upscalingMode: Default
idleTimeoutSeconds: 60
headGroupSpec:
rayStartParams:
num-cpus: '1'
node-manager-port: '6380'
object-manager-port: '6381'
runtime-env-agent-port: '6382'
dashboard-agent-grpc-port: '6383'
dashboard-agent-listen-port: '52365'
metrics-export-port: '8080'
max-worker-port: '10012'
# kubeflow-user-example-com should be replaced with the namespace where the Ray cluster is being deployed
# TODO automatically use the current namespace
node-ip-address: $(hostname -I | tr -d ' ' | sed 's/\./-/g').raycluster-istio-headless-svc.kubeflow-user-example-com.svc.cluster.local
template:
metadata:
labels:
sidecar.istio.io/inject: "true"
spec:
serviceAccountName: default-editor
containers:
- name: ray-head # TODO why call it headless with a head...
image: rayproject/ray:2.44.1-py311-cpu
lifecycle:
preStop:
exec:
command: ["/bin/sh","-c","ray stop"]
volumeMounts:
- mountPath: /tmp/ray
name: ray-logs
resources:
limits:
cpu: "1"
memory: "2G"
requests:
cpu: "100m"
memory: "2G"
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
volumes:
- name: ray-logs
emptyDir: {}
workerGroupSpecs:
- replicas: 1
minReplicas: 1
maxReplicas: 1
groupName: small-group
rayStartParams:
num-cpus: '1'
node-manager-port: '6380'
object-manager-port: '6381'
runtime-env-agent-port: '6382'
dashboard-agent-grpc-port: '6383'
dashboard-agent-listen-port: '52365'
metrics-export-port: '8080'
max-worker-port: '10012'
# kubeflow-user-example-com should be replaced with the namespace where the Ray cluster is being deployed
# TODO automatically use the current namespace
node-ip-address: $(hostname -I | tr -d ' ' | sed 's/\./-/g').raycluster-istio-headless-svc.kubeflow-user-example-com.svc.cluster.local
template:
metadata:
labels:
sidecar.istio.io/inject: "true"
spec:
serviceAccountName: default-editor
containers:
- name: ray-worker
image: rayproject/ray:2.44.1-py311-cpu
lifecycle:
preStop:
exec:
command: ["/bin/sh","-c","ray stop"]
# use volumeMounts.Optional.
# Refer to https://kubernetes.io/docs/concepts/storage/volumes/
volumeMounts:
- mountPath: /tmp/ray
name: ray-logs
resources:
limits:
cpu: "1"
memory: "1G"
requests:
cpu: "300m"
memory: "1G"
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
volumes:
- name: ray-logs
emptyDir: {}
+90
View File
@@ -0,0 +1,90 @@
#!/bin/bash
set -euxo
NAMESPACE=$1
TIMEOUT=120 # timeout in seconds
SLEEP_INTERVAL=30 # interval between checks in seconds
RAY_VERSION=2.44.1
start_time=$(date +%s)
for ((i=0; i<TIMEOUT; i+=2)); do
if [[ $(kubectl get namespace $NAMESPACE --no-headers 2>/dev/null | wc -l) -eq 1 ]]; then
echo "Namespace $NAMESPACE created."
break
fi
current_time=$(date +%s)
elapsed_time=$((current_time - start_time))
if [ "$elapsed_time" -ge "$TIMEOUT" ]; then
echo "Timeout exceeded. Namespace $NAMESPACE not created."
exit 1
fi
echo "Waiting for namespace $NAMESPACE to be created..."
sleep 2
done
echo "Namespace $NAMESPACE has been created!"
kubectl label namespace $NAMESPACE istio-injection=enabled
kubectl get namespaces --selector=istio-injection=enabled
# Install KubeRay operator
kustomize build kuberay-operator/overlays/kubeflow | kubectl -n kubeflow apply --server-side -f -
# Wait for the operator to be ready.
kubectl -n kubeflow wait --for=condition=available --timeout=600s deploy/kuberay-operator
kubectl -n kubeflow get pod -l app.kubernetes.io/component=kuberay-operator
# Install RayCluster components
kubectl -n $NAMESPACE apply -f raycluster_example.yaml
# Wait for the RayCluster to be ready.
sleep 5
kubectl -n $NAMESPACE wait --for=condition=Ready pod -l ray.io/cluster=kubeflow-raycluster --timeout=900s
kubectl -n $NAMESPACE logs -l ray.io/cluster=kubeflow-raycluster,ray.io/node-type=head
# Forward the port of Dashboard
sleep 5
kubectl -n $NAMESPACE port-forward --address 0.0.0.0 svc/kubeflow-raycluster-head-svc 8265:8265 &
PID=$!
echo "Forward the port 8265 of Ray head in the background process: $PID"
# Send a curl command to test basic Ray functionality.
sleep 5
output=$(curl -H "Content-Type: application/json" localhost:8265/api/version)
echo "output: ${output}"
# output format: {"version": ..., "ray_version": RAY_VERSION, "ray_commit": ...}
if echo "${output}" | grep -q $RAY_VERSION; then
echo "Test succeeded!"
else
echo "Test failed!"
exit 1
fi
# Delete RayCluster
kubectl -n $NAMESPACE delete -f raycluster_example.yaml
# Wait for all Ray Pods to be deleted.
start_time=$(date +%s)
for ((i=0; i<TIMEOUT; i+=SLEEP_INTERVAL)); do
pods=$(kubectl -n $NAMESPACE get pods -o json | jq '.items | length')
if [ "$pods" -eq 0 ]; then
kill $PID
break
fi
current_time=$(date +%s)
elapsed_time=$((current_time - start_time))
if [ "$elapsed_time" -ge "$TIMEOUT" ]; then
echo "Timeout exceeded. Exiting loop."
exit 1
fi
sleep $SLEEP_INTERVAL
done
# Delete KubeRay operator
kustomize build kuberay-operator/base | kubectl -n kubeflow delete -f -