Services
AI Infrastructure Services
Not bookable yet — I'm building this expertise publicly, with real benchmarks, through a structured hands-on program. Follow the blog for the ongoing build; this list becomes a real service line once it's backed by shipped work, not before.
GPU Scheduling & Sharing on Kubernetes
NVIDIA GPU Operator setup, MIG/MPS/vGPU planning, and time-slicing — getting GPU capacity safely and predictably shared across workloads.
LLM Inference Serving & Optimization
Deploying and tuning vLLM on Kubernetes via KServe — quantization (AWQ/GPTQ), batching, and autoscaling on real serving signals like queue depth rather than CPU/GPU utilization.
Multi-Tenant GPU Platform Engineering
Quota, isolation, and per-tenant cost attribution for shared GPU platforms — Kueue, RBAC, and policy enforcement across teams.
AI Infrastructure Observability & Monitoring
GPU metrics dashboards built on DCGM Exporter, Prometheus, and Grafana — utilization, memory, temperature, and per-pod/namespace attribution, so problems are visible before they cause impact.
AI Infrastructure Cost Optimization
Measuring and reducing cost per million tokens through quantization and batching tuning, with before/after numbers to back it.
Kubernetes Services
Available now — based on current certifications and production experience
Kubernetes Platform Architecture & Setup
Design and deployment of Kubernetes clusters (including Azure Kubernetes Service) with Infrastructure-as-Code via Terraform. Also covers taking over an existing platform from another party.
Kubernetes Day-2 Operations & Reliability
Cluster upgrades, incident resolution, and troubleshooting of workloads, networking, and storage — for platforms already running in production.
GitOps & CI/CD
Implementation and management of Argo CD, Helm chart development, and CI/CD pipeline design or migration.
Observability & Monitoring
Setting up Prometheus and Grafana, dashboards and alerting, so problems are visible before they cause impact.
Service Mesh & Networking
Implementation of Istio or Linkerd, ingress controller setup, and node lifecycle automation.
Security Hardening & Compliance
Identifying and remediating vulnerabilities, and preparing the platform for a security audit.
Platform Assessment & Advisory
An independent analysis of your current Kubernetes environment, with concrete recommendations on reliability, maintainability, and operational maturity.
Training & Team Enablement
Kubernetes training and coaching for your own engineering team, aimed at managing the platform independently long-term.
Need Kubernetes expertise on your team or platform?
Get in touch via the contact form, or email directly at contact@albertpruis.com.