GPU Development Setup¶
This guide covers setting up a local development environment with NVIDIA GPU support for running vLLM and kvcached.
Prerequisites¶
Hardware¶
- NVIDIA GPU with CUDA 13.x drivers
- 8GB+ VRAM (16GB+ recommended for larger models)
Software¶
- Node.js 22.x - See main README for installation
- Python 3.12 - Required for vLLM (auto-installed by uv on first run)
- uv - Fast Python package manager
Verify GPU Setup¶
# Check NVIDIA driver and CUDA version
nvidia-smi
# Expected output shows GPU name, driver version, CUDA version
# Example: NVIDIA GeForce RTX 3070, Driver 565.xx, CUDA 13.x
Install uv (Python Package Manager)¶
Quick Start¶
Once prerequisites are installed, the development scripts handle everything automatically:
# Install Node.js dependencies
npm install
npm run build -w packages/types && npm run build -w packages/utils
# Start development (auto-creates Python venv on first run)
npm run dev
On first run, the script will:
- Create a Python virtual environment at
apps/backend/.venv - Install dependencies from
apps/backend/pyproject.toml(vLLM, kvcached) - Set up kvcached environment variables
- Start the backend and frontend development servers
How It Works¶
Environment Variables¶
The backend uses environment variables for configuration. A reference file is available at apps/backend/.env.example.
kvcached variables (configured by start-dev script):
| Variable | Default | Description |
|---|---|---|
ENABLE_KVCACHED |
true |
Enables kvcached memory sharing |
KVCACHED_AUTOPATCH |
1 |
Auto-patches vLLM for kvcached support |
CUDA_VISIBLE_DEVICES |
0 |
GPU device index |
HuggingFace authentication (for gated models like Llama):
| Variable | Default | Description |
|---|---|---|
HF_TOKEN |
(none) | HuggingFace access token for gated models |
Get your token from HuggingFace Settings. You can also set this via the Settings page in the web UI.
Model Loading¶
Models are loaded dynamically via the API or Admin Dashboard - no models are pre-loaded at startup. When you load a model:
- Backend receives the load request
- vLLM process is spawned with kvcached environment
- Model is downloaded (if not cached) and loaded to GPU
- Multiple models can share GPU memory via kvcached
Recommended Models for 8GB VRAM¶
For development and testing on 8GB GPUs:
| Model | Parameters | VRAM Usage | Notes |
|---|---|---|---|
Qwen/Qwen3-0.6B |
0.6B | ~2GB | Ideal for testing, fast inference |
meta-llama/Llama-3.2-1B |
1B | ~3GB | Good balance of size and capability |
microsoft/phi-2 |
2.7B | ~6GB | Near 8GB limit, good quality |
TinyLlama/TinyLlama-1.1B-Chat-v1.0 |
1.1B | ~3GB | Chat-optimized small model |
Memory Tips¶
- Use
--gpu-memory-utilization 0.3for small models to leave room for others - With kvcached, you can load 2-3 small models on an 8GB GPU
- Monitor VRAM with
nvidia-smior the Admin Dashboard
Python Dependency Management¶
Python dependencies are managed via apps/backend/pyproject.toml using uv.
Dependency Groups¶
| Group | Packages | Used In |
|---|---|---|
| (base) | kvcached | Dev + Docker |
| dev | vllm | Dev only (Docker base image has vLLM) |
Adding New Packages¶
- Edit
apps/backend/pyproject.toml: - Add to
dependenciesfor packages needed in both dev and production -
Add to
[dependency-groups] devfor development-only packages -
Regenerate lock file:
- Install updated dependencies:
- Commit both
pyproject.tomlanduv.lock
Manual Python Setup (Optional)¶
If you prefer to set up the Python environment manually:
cd apps/backend
# Create venv and install all dependencies (including dev group)
uv sync --locked --group dev
# Activate the environment
source .venv/bin/activate
Then run the server without the wrapper:
# Set environment variables manually
export ENABLE_KVCACHED=true
export KVCACHED_AUTOPATCH=1
export CUDA_VISIBLE_DEVICES=0
# Run server directly
npm run dev:server -w apps/backend
Virtual GPU Mode (Testing Multi-GPU Features)¶
If you only have a single GPU but need to test multi-GPU features like the model move feature, you can enable virtual GPU mode. This creates multiple virtual GPUs that all map to your single physical GPU.
Usage¶
# Start backend with 2 virtual GPUs
DEV_VIRTUAL_GPU_COUNT=2 npm run dev -w apps/backend
# In another terminal
npm run dev -w apps/frontend
What It Does¶
- GPU Detection: Returns N virtual GPUs based on your real GPU's specs
- Model Assignment: Models are assigned to virtual GPU IDs (0, 1, etc.)
- Physical Mapping: All models actually run on physical GPU 0 via
CUDA_VISIBLE_DEVICES=0 - Memory Display: Each virtual GPU shows the real GPU's memory stats
What You Can Test¶
| Feature | Testable? |
|---|---|
| Model move between GPUs | ✅ |
| Move phase transitions | ✅ |
| Connection draining | ✅ |
| Cancel/rollback semantics | ✅ |
| Multi-GPU UI display | ✅ |
| Actual GPU memory isolation | ❌ |
| Real tensor parallelism | ❌ |
Limitations¶
- All virtual GPUs share the same physical GPU memory
- Memory tracking per virtual GPU is approximate
- Cannot test actual multi-GPU tensor parallelism
- Loading too many models will exhaust real GPU memory
Example Workflow¶
- Start backend with
DEV_VIRTUAL_GPU_COUNT=2 - Open UI - you'll see 2 GPUs listed
- Load a model on vGPU 0
- Use "Move to different GPU" to move it to vGPU 1
- Verify the move completes through all phases
- Model now appears on vGPU 1 in the UI
GPU-Free Development (Inference Simulator)¶
If you don't have a GPU, you can use the inference simulator — a lightweight Go binary (llm-d-inference-sim) that mimics vLLM without requiring GPU hardware or NVML drivers. It provides OpenAI-compatible endpoints, simulated model loading, and realistic timing behavior.
Installing the Simulator Binary¶
Build from source (requires Go 1.22+):
git clone https://github.com/llm-d/llm-d-inference-sim.git
cd llm-d-inference-sim
make build
sudo cp bin/llm-d-inference-sim /usr/local/bin/
Or download a pre-built release from the GitHub releases page.
Verify the installation:
Quick Start¶
A convenience script runs both backend and frontend with simulator defaults:
This uses the preset in apps/backend/.env.sim (4 simulated GPUs, 24 GB each). Edit that file to adjust the configuration.
You can also set variables inline:
# Single GPU with 48 GB
INFERENCE_BACKEND=inference-sim DEV_VIRTUAL_GPU_COUNT=1 SIM_GPU_MEMORY_GB=48 npm run dev
Configuration Reference¶
| Variable | Default | Description |
|---|---|---|
INFERENCE_BACKEND |
vllm |
Set to inference-sim to enable the simulator |
DEV_VIRTUAL_GPU_COUNT |
0 (→ 1 in sim mode) |
Number of simulated GPUs |
SIM_GPU_MEMORY_GB |
24 |
VRAM per simulated GPU (GB) |
SIM_MODEL_MEMORY_GB |
4 |
Default model memory estimate when size can't be detected (GB) |
SIM_STARTUP_DURATION |
3s |
Simulated model loading time (Go duration format) |
INFERENCE_SIM_BINARY |
llm-d-inference-sim |
Path to the binary (if not in PATH) |
How Model Memory Estimation Works¶
The simulator estimates GPU memory usage from the model name:
- Names containing a size like
7Bor1.5B→ estimated asparams × 2GB (fp16) for models ≤30B, orparams × 0.5 + 2GB (quantized) for larger models - Names without a recognizable size (e.g.,
SmolLM2-360M-Instruct) → falls back toSIM_MODEL_MEMORY_GB(default 4 GB)
What Works in Simulator Mode¶
| Feature | Supported |
|---|---|
| Model load / unload | Yes |
| Multi-GPU assignment | Yes |
| Sleep / wake | Yes |
| GPU memory tracking | Yes (simulated) |
| Model move between GPUs | Yes |
| Inference requests | Yes (echo responses) |
| Dashboard & all UI features | Yes |
| Real inference output | No (responses echo the input) |
| Actual GPU memory pressure | No |
| kvcached memory sharing | No |
Multi-Pod Cluster Development¶
You can simulate a multi-pod cluster locally using the inference simulator and CLUSTER_PEERS:
make dev:cluster:sim # 2 pods × 2 GPUs (default)
make dev:cluster:sim PODS=3 GPUS=4 # 3 pods × 4 GPUs
This launches multiple backend instances on sequential ports (3001, 3002, ...) plus the frontend.
Port isolation: vLLM serving ports are automatically partitioned per pod to avoid collisions on localhost. Each pod gets a range of 100 ports based on its position in the CLUSTER_PEERS list:
| Pod | API Port | vLLM Port Range |
|---|---|---|
| 0 | 3001 | 12346–12445 |
| 1 | 3002 | 12446–12545 |
| 2 | 3003 | 12546–12645 |
This is automatic when CLUSTER_PEERS is set and VLLM_BASE_PORT is not explicitly configured. Set VLLM_BASE_PORT to override.
Frontend-Only Development¶
If you only need the frontend (no backend at all):
Developing with OAuth Authentication¶
If you want to test OAuth authentication locally against a remote OpenShift cluster, follow these steps.
Prerequisites¶
- Access to an OpenShift cluster with OAuth configured
ocCLI logged in with namespace admin permissions- A namespace for sardeenz RBAC resources (e.g.,
sardeenz)
Step 1: Create OAuth Client on OpenShift¶
Register a development OAuth client with localhost callback URL:
oc apply -f - <<EOF
apiVersion: oauth.openshift.io/v1
kind: OAuthClient
metadata:
name: sardeenz-dev
grantMethod: auto
redirectURIs:
- http://localhost:3000/api/auth/callback
secret: dev-secret-change-me
EOF
Note: Change the
secretvalue to something unique for your development environment.
Step 2: Create RBAC Resources on OpenShift¶
Create the namespace and apply the RBAC manifests from the deployment folder:
# Create namespace (if it doesn't exist)
oc new-project sardeenz
# Apply ServiceAccount and RBAC resources
oc apply -f deployment/serviceaccount.yaml -n sardeenz
oc apply -f deployment/rbac.yaml -n sardeenz
This creates:
- The
sardeenzServiceAccount - The
sardeenz-adminandsardeenz-admin-readonlymarker Roles - The
sardeenz-auth-reviewerRole and RoleBinding
Step 3: Create RoleBinding for Your User¶
Grant yourself admin access:
For complete RBAC configuration options (adding other users, groups, or all authenticated users), see the RBAC Setup Guide.
Note for Cluster-Admin Users: If you're logged in as a cluster-admin, you will automatically have access to sardeenz regardless of whether you have a sardeenz-admin RoleBinding. This is expected Kubernetes RBAC behavior—cluster-admin has wildcard permissions on all resources including sardeenz marker roles. To properly test RBAC authorization (denied access, read-only vs admin), use a non-cluster-admin account.
Step 4: Generate ServiceAccount Token¶
Generate a long-lived token for local development:
Save this token - you'll need it for the environment variable.
Step 5: Configure Environment Variables¶
Create or update apps/backend/.env with the OAuth configuration:
# Authentication mode
AUTH_MODE=oauth
# OAuth client configuration (must match the OAuthClient created above)
OAUTH_CLIENT_ID=sardeenz-dev
OAUTH_CLIENT_SECRET=dev-secret-change-me
# OpenShift URLs
OAUTH_ISSUER_URL=https://oauth-openshift.apps.your-cluster.com
K8S_API_URL=https://api.your-cluster.com:6443
# RBAC configuration
NAMESPACE=sardeenz
SERVICE_ACCOUNT_TOKEN=<paste token from Step 4>
# Local development URLs
API_BASE_URL=http://localhost:3000
FRONTEND_URL=http://localhost:5173
# JWT secret (any random string for development)
JWT_SECRET=your-dev-jwt-secret-change-me
Getting the OpenShift URLs¶
To find the correct URLs for your cluster:
# Get the Kubernetes API URL
oc whoami --show-server
# Example output: https://api.your-cluster.com:6443
# The OAuth issuer URL is typically derived from the API URL
# Replace 'api.' with 'oauth-openshift.apps.' and remove the port
# Example: https://oauth-openshift.apps.your-cluster.com
Step 6: Start Development¶
Navigate to http://localhost:5173. When you click "Login", you'll be redirected to OpenShift OAuth, then back to the local frontend with your session authenticated.
Troubleshooting OAuth Development¶
"Access denied" after login:
- Verify your RoleBinding exists:
oc get rolebindings -n sardeenz - Check if you have the right role:
oc auth can-i get admin.sardeenz.rh-aiservices-bu.io -n sardeenz --as=$(oc whoami)
"LocalSubjectAccessReview failed" errors:
- Check that the
sardeenz-auth-reviewerRoleBinding exists - Verify your SERVICE_ACCOUNT_TOKEN is valid and not expired
- Ensure K8S_API_URL is correct and reachable from your machine
"Invalid redirect URI" from OpenShift:
- Verify the OAuthClient has
http://localhost:3000/api/auth/callbackin redirectURIs - Ensure API_BASE_URL matches exactly
Troubleshooting¶
"uv is not installed"¶
Install uv:
"nvidia-smi: command not found"¶
NVIDIA drivers are not installed. Install them from:
- Ubuntu:
sudo apt install nvidia-driver-535 - Fedora:
sudo dnf install akmod-nvidia - Or download from NVIDIA's website
"CUDA out of memory"¶
- Use smaller models (see recommendations above)
- Reduce
--gpu-memory-utilizationin model config - Unload unused models via API/Dashboard
- Check for other GPU processes:
nvidia-smi
"vLLM not found" or spawn errors¶
Ensure the virtual environment is activated and vLLM is installed:
cd apps/backend
uv sync --locked --group dev
source .venv/bin/activate
which vllm # Should show path in .venv/bin/
kvcached IPC errors¶
If you see IPC segment errors:
kvcached Tools¶
kvcached provides CLI tools for monitoring:
# Activate the venv first
source apps/backend/.venv/bin/activate
# List active memory segments
kvctl list
# Monitor memory usage in real-time
kvtop
# Set memory limit for a model
kvctl limit <segment-name> 4G
Additional Resources¶
- Architecture Documentation - System design and vLLM integration
- API Guide - Model loading API endpoints
- RBAC Setup Guide - Kubernetes-native RBAC configuration for OAuth
- kvcached Documentation - Detailed kvcached configuration