Posts Tagged ‘Kubernetes’
[GopherConUK2025] Conceptual XPDB Custom Resource definition
apiVersion: policy.form3.tech/v1alpha1
kind: CrossClusterPodDisruptionBudget
metadata:
name: cockroachdb-global-pdb
spec:
maxUnavailable: 1
clusters:
– aws-region-1
– gcp-region-1
– azure-region-1
selector:
matchLabels:
app: cockroachdb
XPDB enforces global disruption caps across cluster boundaries. When node drains or pod evasions occur in one cloud provider, XPDB evaluates the global state across all environments. By guaranteeing that only a single CockroachDB pod is disrupted across the entire global infrastructure at any given moment, XPDB ensures data consensus remains fully protected during routine infrastructure maintenance.
## Operator-Driven Infrastructure and Rolling Node-Pool Management
Managing multi-tenant infrastructure across multiple jurisdictions—each containing distinct development, staging, and production environments—results in a massive expansion of node pools. Updating Kubernetes worker nodes across this matrix using traditional Infrastructure-as-Code tools like Terraform creates extreme operational friction, requiring dozens of sequential pull requests and complex deployment pipelines.
To address this management overhead, Form3 built a custom Kubernetes operator called the Cluster Lifecycle Operator. The operator runs natively inside each managed cluster and abstracts raw node-pool management behind Custom Resource Definitions (CRDs).
// Conceptual snippet of CRD controller reconciliation loop
package main
import (
“context”
“fmt”
)
type ClusterSpec struct {
Version string json:"version"
}
func ReconcileNodePool(ctx context.Context, spec ClusterSpec) error {
fmt.Printf(“Reconciling node pools to version: %s\n”, spec.Version)
// Operator logic handles sequential node drains
return nil
}
Instead of modifying declarative infrastructure files for every individual node pool across every provider, engineers update a single CRD spec controlling the target cluster version. The Cluster Lifecycle Operator handles the rolling replacement of worker nodes asynchronously, adhering to defined disruption parameters and health checks. Upgrading the global fleet requires only three sequential pull requests—promoted systematically through development, staging, and production.
## Continuous Disaster Recovery via Automated Production Chaos Injection
Conventional disaster recovery (DR) practices often rely on periodic manual failover tests driven by static documentation. In rapidly changing microservice environments, compliance-focused DR exercises fail to validate real-world resilience, as system changes can render manual playbooks obsolete immediately after testing.
Form3 enforces continuous disaster recovery validation directly within staging environments through automated fault injection. A dedicated test harness deploys synthetic client applications outside the primary infrastructure perimeter. These synthetic actors continuously execute end-to-end payment workflows against simulated payment scheme interfaces at fixed intervals.
// Custom chaos injection test runner in Go
package main
import (
“context”
“log”
“time”
)
func InjectProviderOutage(ctx context.Context, targetCloud string) error {
log.Printf(“Simulating complete network partition for provider: %s”, targetCloud)
// Inject network block rules via Chaos Mesh API
time.Sleep(30 * time.Second)
return nil
}
“`
During automated test windows, the platform uses Chaos Mesh alongside custom Go-based orchestrators to introduce disruptive scenarios:
- Severing inter-cloud network connectivity.
- Terminating entire database nodes or message brokers.
- Programmatically isolating an entire cloud provider for 24 hours.
If synthetic payment processing encounters errors or breaches latency thresholds, automated alerts notify on-call platform teams. Automated daily reports detail platform behavior during fault injection cycles, verifying that the loss of an entire cloud vendor produces zero client impact.
Links:
[KCDUK2024] Kubernetes Privilege Escalation Tactics: Unveiling Vulnerabilities and Fortifying Defenses
Iain Smart and Andrew Martin presented a compelling exploration of Kubernetes privilege escalation tactics, offering a deep dive into how both trusted and unprivileged users can exploit vulnerabilities within the system. Their discussion, a highlight of KCDUK2024, provided invaluable insights for SREs, security teams, and pentesters aiming to enhance cluster security.
The core of their presentation focused on the critical need to understand potential attack vectors and implement robust defense mechanisms. They articulated that while penetration testing Kubernetes should inherently be challenging, certain oversights can inadvertently simplify the process for malicious actors. The speakers emphasized the multifaceted nature of threats, ranging from rogue SREs and disaffected platform developers to external hostile internet citizens.
Smart and Martin meticulously guided attendees through various methods to escalate privileges, achieve persistence, and potentially wreak havoc across a cluster, all while attempting to obscure any traces of activity. Their expertise illuminated the intricate interplay of Kubernetes components and how unusual interactions or component abuse can be leveraged for unauthorized access.
Understanding Kubernetes Vulnerabilities and Exploitation
The presenters underscored the importance of comprehending the array of Kubernetes vulnerabilities that security professionals must be aware of. They elaborated on specific techniques that adversaries might employ, detailing how seemingly minor misconfigurations or overlooked edge cases can become critical points of entry. The discussion extended to identifying different adversary levels, stressing that tailoring defenses according to the threat model is paramount for effective security.
Smart and Martin provided practical insights into the most cost-effective and efficient strategies for fortifying Kubernetes clusters. They advocated for a proactive approach that encompasses preventative controls, such as static analysis on deployed artifacts and meticulous enumeration of Role-Based Access Control (RBAC). The emphasis was not solely on prevention but also on robust detective controls, including monitoring external traffic through split-horizon DNS to identify suspicious outbound communications.
Mitigating Risks and Ensuring Robust Security
A key takeaway from Smart and Martin’s presentation was the critical role of remediative controls. These controls are designed to detect ongoing attacks and initiate automated responses, such as node draining and shutdown procedures, to prevent data exfiltration. Despite implementing a comprehensive suite of preventative, detective, and remediative measures, they acknowledged that complete isolation and absolute security are unattainable ideals. The ever-evolving threat landscape, exemplified by nation-state attacks and sophisticated persistent threats, necessitates continuous vigilance and adaptation.
The speakers concluded by reiterating that detection is as crucial as prevention in the complex puzzle of cloud-native security. They highlighted that while various security tools and practices are available, the dynamic nature of threats requires an ongoing commitment to learning, monitoring, and adapting. Their insights provided a valuable roadmap for organizations striving to secure their Kubernetes environments against an increasingly sophisticated array of threats.
Links:
[KCDUK2024] The Operator Antipattern | Gerald Schmidt
At KCDUK2024, Gerald Schmidt, data platform engineering lead at DS Smith, delivered a candid reflection on the Kubernetes operator pattern, questioning its efficacy in modern platform engineering. Drawing from his experience at Babylon Health, Go City, and Thoughtworks, Gerald shared a journey from enthusiasm to disillusionment, highlighting how operators, despite their promise, often introduce complexity and fragility that outweigh their benefits.
The Allure and Pitfalls of Operators
Gerald opened with his initial excitement for operators, which promised to extend Kubernetes’ API for managing complex stateful applications. He referenced Brandon Phillips’ vision of operators embedding domain expertise to automate tasks, such as managing databases or message queues. However, Gerald’s experience revealed that operators rarely delivered on this promise. Instead, they introduced tight coupling between custom resource definitions (CRDs) and controllers, creating operational challenges.
A pivotal anecdote from his time at a health startup illustrated this. Gerald’s team built operators for data flows, only to face disaster when CRDs were inadvertently deleted, requiring a weekend-long recovery. This incident led to their team losing deployment privileges, underscoring the risks of operators in environments with shrinking teams and heightened scrutiny. Gerald argued that operators, while powerful, often fail to match the robustness of managed services like AWS RDS, which offer superior durability and availability.
Persistent Volumes Over Stateful Applications
Gerald challenged the notion that Kubernetes struggles with stateful applications, suggesting the real issue lies with persistent volumes. Solutions like Thanos and WarpStream, which leverage object storage, address this effectively without relying on operators. He highlighted the maturity of the object storage ecosystem, with providers like AWS S3, Google Cloud Storage, and Wasabi offering cost-effective, reliable alternatives. The Container Object Storage Interface (COSI), though still emerging, promises further simplification, reducing the need for complex operator-based solutions.
He contrasted this with the operator pattern’s limitations, particularly around CRDs. Unlike controllers, which integrate seamlessly with Kubernetes’ control loop, CRDs require administrative privileges and complicate versioning and upgrades. Gerald cited the AWS Controllers for Kubernetes (ACK), still in v1alpha1, as an example of versioning challenges, where upgrades introduce new failure modes and complexity.
Toward Simpler, Developer-Friendly Solutions
Reflecting on developer experience, Gerald advocated for alternatives like config maps over CRDs. He praised Grafana’s approach, which uses config maps for dashboards, avoiding the need for administrative access and simplifying deployment. Similarly, older Prometheus setups allowed developers to control scraping with simple annotations, granting autonomy without the overhead of CRDs. Gerald argued that controllers alone, paired with config maps, provide most of the operator’s benefits without the downsides.
He concluded with a decision framework: only pursue the operator pattern if the use case justifies a domain-specific language. For operational tasks with low blast radius, like Kyverno’s policy enforcement, operators excel. However, for development teams, Gerald recommended avoiding custom operators, as they often lead to wasted effort and fragility, as seen in his team’s abandoned six-month project.
Links:
[KCDUK2024] CVEs and Kubernetes: A Love Story? | Marcus Tenorio
In a lively lightning talk at KCDUK2024, Marcus Tenorio, an engineering manager with a background in incident response, brought a fresh perspective on the relationship between Kubernetes and Common Vulnerabilities and Exposures (CVEs). With a nod to the conference’s community spirit, Marcus framed security challenges as opportunities for growth, likening the evolution of Kubernetes security to a love story where vulnerabilities drive collaboration and improvement.
The Evolution of Kubernetes Security
Marcus began by exploring the CVE landscape, drawing from the official CVE feed, MITRE, and NVD databases. He noted that while Kubernetes, launched in 2014, has seen a rise in reported CVEs, this reflects increased scrutiny rather than declining security. Early vulnerabilities, often identified by community members like a Google engineer on GitHub, showcased the power of open-source collaboration. Marcus highlighted that critical CVEs in Kubernetes are relatively rare, contrasting with infamous incidents like Log4j, suggesting a stable core.
He analyzed a sample of 55 CVEs, revealing that the growth in reported vulnerabilities corresponds to Kubernetes’ maturity. As the platform evolves, the community actively identifies and resolves issues, strengthening its security posture. Marcus emphasized that this process mirrors a relationship where challenges foster growth, with each CVE contributing to a more robust ecosystem.
Community-Driven Security
The heart of Marcus’s talk was the role of community in Kubernetes security. He shared an anecdote from an e-commerce platform where a team’s proactive vulnerability hunting led to safer systems, not because vulnerabilities were abundant, but because they were addressed collaboratively. This approach, rooted in policies and shared learning, transforms potential threats into opportunities for improvement.
Marcus encouraged attendees to embrace this “love” for security by fostering open communication and leveraging data to understand vulnerabilities. Tools like Datadog, despite occasional AI hallucinations, help teams analyze and respond to CVEs effectively. By viewing security as a collective journey, Marcus underscored how Kubernetes’ community-driven model drives resilience, aligning with KCDUK2024’s ethos of collaboration.
[KCDUK2024] Platform Orchestrators: The Missing Middle of Internal Developer Platforms | Daniel Bryant
At KCDUK2024, Daniel Bryant, a seasoned platform engineering advocate, delivered a compelling case for platform orchestrators as the critical “missing middle” in internal developer platforms (IDPs). Drawing from his extensive experience with Kubernetes, Mesos, and tools like Backstage and Crossplane, Daniel explored how orchestrators bridge the gap between developer-facing portals and infrastructure layers, enabling scalable, efficient, and user-centric platforms. His talk offered a blueprint for organizations to balance speed, safety, and scale in their platform engineering efforts.
The Evolution of Platform Engineering
Daniel began by contrasting three approaches to platform building: top-down, app-centric portals; bottom-up, infrastructure-focused solutions; and a middle-out, platform-engineering-focused model. Top-down approaches, like Backstage, excel at providing quick wins with developer portals but struggle with day-two operations like upgrades and maintenance. Bottom-up approaches, such as Terraform or Crossplane, offer robust automation but often overwhelm developers with infrastructure complexity. The middle-out approach, which Daniel champions, treats platforms as products, prioritizing user needs and process automation.
He referenced Gartner’s platform engineering model, which identifies three layers: application choreography, platform orchestration, and infrastructure composition. The platform orchestration layer, often overlooked, manages the platform’s lifecycle and APIs, ensuring seamless integration between developer workflows and infrastructure. Daniel’s experience with tools like Crossplane and CNOE (Cloud Native Operational Excellence) highlighted how orchestrators codify business processes, reducing coordination overhead and enabling scalability.
Addressing Developer Pain Points
Modern software engineering faces challenges like slow delivery, high-risk deployments, and tech sprawl. Daniel cited statistics showing that 50% of organizations deploy code less than once a month, and 42% of developers fear production failures. Platform orchestrators address these by offering “everything as a service,” from databases to domain-specific services like fraud detection in finance. By automating manual processes, such as security sign-offs, orchestrators enhance safety and efficiency, allowing developers to focus on coding.
Daniel emphasized the importance of progressive disclosure—presenting simple interfaces initially while enabling advanced functionality as needed. He recounted a past experience where a 500-line YAML configuration overwhelmed developers, underscoring the need for intuitive abstractions. Tools like Open Application Model (OAM) and Score, donated to the CNCF, provide developer-friendly APIs, while orchestrators like Kritik and Cusion Stack manage complex workflows, ensuring platforms remain adaptable to changing business needs.
Building Platforms as Products
The heart of Daniel’s message was treating platforms as products, designed with user needs at the forefront. He advocated for clear domain boundaries, inspired by principles like SOLID and CUPID, to ensure platforms are composable and maintainable. Tools like Kritik, where Daniel contributes, use Kubernetes CRDs to define platform APIs, allowing workflows to be containerized and reusable. This approach enables teams to manage platform components at scale, from rolling out security fixes to integrating auditing processes.
Drawing from Team Topologies, Daniel stressed collaboration between platform and development teams to align on goals like adoption rates and onboarding times. He warned against the “build it and they will come” mentality, urging platform engineers to engage users early and measure success through leading indicators like onboarding efficiency and lagging indicators like incident reduction. By treating platforms as products, organizations can achieve the speed, safety, and scale needed to thrive in cloud-native environments.
Links:
Understanding Kubernetes for Docker and Docker Compose Users
TL;DR
Kubernetes may look like an overly complicated version of Docker Compose, but it operates on a different level entirely. Where Compose excels at quick, local orchestration of containers, Kubernetes is a robust, distributed platform designed for automated scaling, fault-tolerance, and production-grade deployments across multi-node clusters. This article provides a comprehensive comparison and shows how ArgoCD enhances GitOps-based Kubernetes workflows.
Docker Compose vs Kubernetes – Similarities and First Impressions
At a high level, Docker Compose and Kubernetes share similar concepts: containers, services, configuration, and volumes. This often leads to the assumption that Kubernetes is just a verbose, harder-to-write Compose replacement. However, Kubernetes is more than a runtime. It’s a control plane, a state manager, and a policy enforcer.
| Concept | Docker Compose | Kubernetes |
|---|---|---|
| Service definition | docker-compose.yml |
Deployment, Service, etc. YAML manifests |
| Networking | Shared bridge network, service discovery by name | DNS, internal IPs, ClusterIP, NodePort, Ingress |
| Volume management | volumes: |
PersistentVolume, PersistentVolumeClaim, StorageClass |
| Secrets and configs | .env, environment: |
ConfigMap, Secret, ServiceAccount |
| Dependency management | depends_on |
initContainers, readinessProbe, livenessProbe |
| Scaling | Manual (scale flag or duplicate services) | Declarative (replicas), automatic via HPA |
Real-Life Use Cases – Docker Compose vs Kubernetes Examples
Tomcat + Oracle + MongoDB + NGINX Stack
Docker Compose
version: '3'
services:
nginx:
image: nginx:latest
ports:
- "80:80"
depends_on:
- tomcat
tomcat:
image: tomcat:9
ports:
- "8080:8080"
environment:
DB_URL: jdbc:oracle:thin:@oracle:1521:orcl
oracle:
image: oracle/database:19.3.0-ee
environment:
ORACLE_PWD: secretpass
volumes:
- oracle-data:/opt/oracle/oradata
mongo:
image: mongo:5
volumes:
- mongo-data:/data/db
volumes:
oracle-data:
mongo-data:
Kubernetes Equivalent
- Each service becomes a
Deploymentand aService. - Environment variables and passwords are stored in
Secrets. - Volumes are defined with
PVCandStorageClass.
apiVersion: v1
kind: Secret
metadata:
name: oracle-secret
type: Opaque
data:
ORACLE_PWD: c2VjcmV0cGFzcw==
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: tomcat
spec:
replicas: 2
selector:
matchLabels:
app: tomcat
template:
metadata:
labels:
app: tomcat
spec:
containers:
- name: tomcat
image: tomcat:9
ports:
- containerPort: 8080
env:
- name: DB_URL
value: jdbc:oracle:thin:@oracle:1521:orcl
NodeJS + Express + MySQL + NGINX
Docker Compose
services:
mysql:
image: mysql:8
environment:
MYSQL_ROOT_PASSWORD: rootpass
volumes:
- mysql-data:/var/lib/mysql
api:
build: ./api
environment:
DB_USER: root
DB_PASS: rootpass
DB_HOST: mysql
nginx:
image: nginx:latest
ports:
- "80:80"
Kubernetes Equivalent
apiVersion: v1
kind: Secret
metadata:
name: mysql-secret
type: Opaque
data:
MYSQL_ROOT_PASSWORD: cm9vdHBhc3M=
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: api
spec:
replicas: 2
template:
spec:
containers:
- name: api
image: node-app:latest
env:
- name: DB_PASS
valueFrom:
secretKeyRef:
name: mysql-secret
key: MYSQL_ROOT_PASSWORD
⚙️ Docker Compose vs kubectl – Command Mapping
| Task | Docker Compose | Kubernetes |
|---|---|---|
| Start services | docker-compose up -d |
kubectl apply -f . |
| Stop/cleanup | docker-compose down |
kubectl delete -f . |
| View logs | docker-compose logs -f |
kubectl logs -f pod-name |
| Scale a service | docker-compose up --scale web=3 |
kubectl scale deployment web --replicas=3 |
| Shell into container | docker-compose exec app sh |
kubectl exec -it pod-name -- /bin/sh |
ArgoCD – GitOps Made Practical
ArgoCD is a Kubernetes-native continuous deployment tool. It uses Git as the single source of truth, enabling declarative infrastructure and GitOps workflows.
✨ Key Features
- Declarative sync of Git and cluster state
- Drift detection and automatic repair
- Multi-environment and multi-namespace support
- CLI and Web UI available
Example ArgoCD Commands
argocd login argocd.myorg.com
argocd app create my-app \
--repo https://github.com/org/app.git \
--path k8s \
--dest-server https://kubernetes.default.svc \
--dest-namespace production
argocd app sync my-app
argocd app get my-app
argocd app diff my-app
Sample ArgoCD Application Manifest
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: my-api
spec:
destination:
namespace: default
server: https://kubernetes.default.svc
project: default
source:
path: k8s/app
repoURL: https://github.com/org/api.git
targetRevision: HEAD
syncPolicy:
automated:
prune: true
selfHeal: true
✅ Conclusion
Docker Compose is perfect for prototyping and local dev. Kubernetes is built for cloud-native workloads, distributed systems, and high availability. ArgoCD makes declarative, Git-based continuous deployment simple, scalable, and observable.
[DevoxxUK2025] Platform Engineering: Shaping the Future of Software Delivery
Paula Kennedy, co-founder and COO of Cintaso, delivered a compelling lightning talk at DevoxxUK2025, tracing the evolution of platform engineering and its impact on software delivery. Drawing from over a decade of experience, Paula explored how platforms have shifted from siloed operations to force multipliers for developer productivity. Referencing the journey from DevOps to PaaS to Kubernetes, she highlighted current trends like inner sourcing and offered practical strategies for assessing platform maturity. Her narrative, infused with lessons from the past and present, underscored the importance of a user-centered approach to avoid the pitfalls of hype and ensure platforms drive innovation.
The Evolution of Platforms
Paula began by framing platforms as foundations that elevate development, drawing on Gregor Hohpe’s analogy of a Volkswagen chassis enabling diverse car models. She recounted her career, starting in 2002 at Acturus, a SaaS provider with rigid silos between developers and operations. The DevOps movement, sparked in 2009, sought to bridge these divides, but its “you build it, you run it” mantra often overwhelmed teams. The rise of Platform-as-a-Service (PaaS), exemplified by Cloud Foundry, simplified infrastructure management, allowing developers to focus on code. However, Paula noted, the complexity of Kubernetes led organizations to build custom internal platforms, sometimes losing sight of the original value proposition.
Current Trends and Challenges
Today, platform engineering is at a crossroads, with Gartner predicting that by 2026, 80% of large organizations will have dedicated teams. Paula highlighted principles like self-service APIs, internal developer portals (e.g., Backstage), and golden paths that guide developers to best practices. She emphasized treating platforms as products, applying product management practices to align with user needs. However, the 2024 DORA report reveals challenges: while platforms boost organizational performance, they often fail to improve software reliability or delivery throughput. Paula attributed this to automation complacency and “platform complacency,” where trust in internal platforms leads to reduced scrutiny, urging teams to prioritize observability and guardrails.
Links:
[KCDUK2024] The Joy of DevEx: Tightening Developer Feedback Loops in Kubernetes | Matthew Revell-Gordon
At KCDUK2024, Matthew Revell-Gordon, a senior platform engineering consultant at OpenCredo, shared an engaging exploration of how to enhance the developer experience (DevEx) by tightening feedback loops in Kubernetes. His talk addressed the challenges of local testing for cloud-native applications, offering practical solutions to align local Kubernetes environments with cloud deployments. By leveraging tools like Colima, KIND, and MetalLB, Matthew demonstrated how to create a robust local setup that mirrors production, fostering faster development cycles and a more joyful DevEx.
Challenges of Local Testing in Cloud-Native Environments
Cloud-native technologies like Kubernetes are designed to thrive in distributed, scalable cloud environments, which poses a significant hurdle for local testing. Matthew highlighted that while unit tests and linting are straightforward on a laptop, integration testing and infrastructure configuration validation are far more complex. Laptops lack the scale and fidelity of cloud environments, leading developers to rely heavily on cloud-based testing. This approach, however, introduces lengthy feedback loops, as CI/CD pipelines can take time to queue, bootstrap, and execute—often failing due to trivial errors like typos that could have been caught earlier.
Matthew recounted a real-world scenario where developers used Docker Compose for local testing, only to encounter issues when deploying to Kubernetes due to untested Ingress controller changes. This mismatch delayed releases and frustrated teams, underscoring the need for a local environment that closely emulates production. He emphasized that while cloud testing is inevitable, optimizing local setups can significantly reduce cycle times, catching errors before they reach costly cloud pipelines.
Crafting a Production-Like Local Kubernetes Setup
To address these challenges, Matthew proposed a solution centered on Colima, KIND, and MetalLB to create a local Kubernetes cluster that mirrors cloud deployments. Colima, a lightweight alternative to Docker Desktop, provides a configurable Linux-based virtual machine, allowing developers to fine-tune their environment. KIND (Kubernetes IN Docker) enables multi-node, high-availability clusters with precise control over Kubernetes versions, ensuring consistency with production. MetalLB replaces cloud-specific load balancer controllers, provisioning external IP addresses for local services.
In a live demo, Matthew provisioned a multi-node Kubernetes cluster using KIND, deploying a simple application with node affinity and a load balancer service accessible directly from his Mac. He addressed networking challenges by configuring routes and IP tables, ensuring seamless connectivity between the host, Colima VM, and KIND network. This setup, while not identical to a cloud environment, offers a close approximation, enabling developers to test complex configurations like node affinity and rollouts locally. Matthew also recommended automating these setups with scripts, making them accessible to developers who may not be infrastructure experts, thus enhancing team efficiency.
Enhancing DevEx Through Automation and Community
The true joy of DevEx, Matthew argued, lies in empowering developers to focus on coding rather than wrestling with infrastructure. By packaging local Kubernetes setups into one-click scripts, platform engineers can democratize access to production-like environments, reducing friction and boosting productivity. He cautioned against anti-patterns like creating local-only manifests or altering production setups for testing, as these diverge from the goal of fidelity. Instead, he advocated an iterative approach, starting with the most pressing pain points and gradually refining the local environment.
Matthew’s passion for community resonated throughout his talk. He encouraged platform engineers to share their tooling and scripts, fostering collaboration and easing the burden on developers. By tightening feedback loops, these solutions not only save time and costs but also cultivate a sense of joy in the development process, aligning with the broader ethos of KCDUK2024’s focus on community-driven innovation.
Links:
[KCDUK2024] Comprehensible Kubernetes: Empowering Scientists with Scalable and Secure Platforms for HPC and AI
At KCDUK2024, Scott Coulton and Tyler, representing StackHPC, delivered an insightful presentation on making Kubernetes accessible to scientists and researchers with minimal technical backgrounds. Their talk focused on crafting cloud-native platforms that prioritize usability, security, and scalability, enabling researchers to focus on their work rather than grappling with complex infrastructure. By leveraging open-source tools like ClusterAPI, Helm, Zenith, and Keycloak, StackHPC has bridged the gap between high-performance computing (HPC) and cloud-native technologies, transforming how scientific research is conducted.
Bridging HPC and Cloud-Native for Research
StackHPC, a Bristol-based consultancy specializing in HPC and cloud solutions, was founded on traditional HPC expertise, such as Slurm clusters and high-performance networking. Scott and Tyler outlined their mission to translate this expertise into the cloud-native realm, primarily for universities and research institutions. Their approach centers on three pillars: reconfigurable infrastructure, performance optimization, and self-service applications. By deploying private OpenStack clouds and Kubernetes clusters, they enable researchers to access tailored environments without needing deep technical knowledge.
The diversity of scientific use cases presents unique challenges. Researchers may require Slurm clusters for batch processing, JupyterHub for data analysis, or GPU-enabled environments for machine learning. Scott highlighted a common thread: scientists are not platform engineers and seek intuitive, reliable platforms. StackHPC’s solution, the LOKI stack (Linux, OpenStack, Kubernetes Infrastructure), integrates open-source tools to provide a flexible, self-service platform that meets these needs while maintaining security and scalability.
The LOKI Stack: A Scalable Solution
Tyler delved into the technical underpinnings of StackHPC’s LOKI stack, which combines ClusterAPI, Zenith, and Azimuth to deliver seamless infrastructure management. ClusterAPI, a declarative API, simplifies Kubernetes cluster lifecycle management, supporting auto-healing and auto-scaling across providers like OpenStack. Tyler explained how Helm charts streamline cluster provisioning, ensuring consistency and ease of deployment. This approach allows operators to manage infrastructure efficiently, freeing researchers from administrative burdens.
Zenith, an innovative application proxy, enables secure external access to applications without public IPs, using SSH tunneling and OIDC authentication. Azimuth, a self-service web portal, offers pre-configured appliances like JupyterHub and Slurm clusters, customizable via Helm charts. Keycloak integration ensures secure access management, allowing platform administrators to create user accounts without granting cloud access. This architecture empowers researchers to deploy and manage platforms independently, aligning complexity with their expertise.
Case Studies: Real-World Impact
Scott presented two case studies illustrating the LOKI stack’s versatility. The first involved deploying JupyterHub for training courses on a private OpenStack cloud. Using Azimuth, administrators created isolated Keycloak realms to manage attendee accounts, ensuring secure access via Zenith’s proxying. This setup allowed course tutors to focus on teaching, with attendees accessing familiar JupyterHub environments without needing cloud credentials. The solution’s simplicity and security made it ideal for educational settings.
The second case study addressed a university’s request for a privacy-preserving large language model (LLM) service. StackHPC developed a Helm chart to deploy open-source LLMs, integrated with Azimuth for ad-hoc testing and ArgoCD for production-grade management. The service, designed to be GDPR-compliant, provided an anonymous interface for staff and students to experiment with LLMs. Monitoring via Prometheus and Grafana ensured reliability, demonstrating how StackHPC’s stack adapts to emerging technologies while maintaining robustness.
Empowering Researchers Through Simplicity
The core takeaway from Scott and Tyler’s talk was that scientists prioritize research over infrastructure management. By offering a platform that balances infrastructure-as-a-service flexibility with platform-as-a-service simplicity, StackHPC empowers researchers to work efficiently. Their commitment to open-source, evidenced by plans to donate Zenith and Azimuth to the CNCF sandbox, underscores their dedication to community-driven innovation. The LOKI stack’s ability to abstract complexity while preserving functionality positions it as a transformative tool for scientific computing.
Links:
[KCDUK2024] From Free Kicks to Git Commits: Steve Wade’s Journey at KCDUK2024
At KCDUK2024, Steve Wade, a cloud native consultant and trainer at Jetstack, captivated the audience with a narrative of transformation, tracing his remarkable journey from professional football to a thriving career in technology. His talk, a blend of personal reflection and technical insight, illuminated the parallels between orchestrating a football team and building self-service platforms on Kubernetes. Steve’s story is one of resilience, adaptability, and the power of transferable skills, offering a compelling blueprint for navigating career pivots and leveraging past experiences to excel in the tech landscape.
From the Pitch to the Keyboard: A Career Pivot
Steve’s journey began at the tender age of four, scouted for his football potential, a moment that ignited a lifelong passion. His early career was defined by dreams of playing in iconic stadiums like Wembley or in a Champions League final. However, a devastating knee injury at 18—shattering his ACL and ligaments—halted his aspirations, plunging him into an identity crisis. Unable to walk for 18 months, Steve faced a profound setback, spiraling into depression and grappling with the question, “Who am I if not a footballer?” This period of introspection, however, became a catalyst for reinvention.
Inspired by his father’s work in IT, Steve’s curiosity was piqued. From his bed, he began exploring technology, diving into coding and the intricacies of programming languages. This marked the beginning of a significant pivot, transitioning from the physical demands of the football field to the intellectual challenges of the digital realm. His relentless drive, honed on the pitch, translated into a disciplined approach to learning, setting the foundation for his eventual expertise in cloud native technologies.
Lessons from Football: Teamwork, Empathy, and Resilience
Steve eloquently drew parallels between the skills cultivated in football and those essential in technology. On the pitch, teamwork was paramount—11 players working cohesively toward a single goal, with intricate passes leading to success. Similarly, in tech, delivering a product requires collaboration across diverse teams, where no one “scores” alone. Steve emphasized empathy as a cornerstone, recounting how supporting teammates in football translated to understanding developers’ needs in tech. Knowing when a colleague needs support, he argued, transforms good teams into great ones, a principle that resonates in both domains.
Resilience, another lesson from his sporting days, proved invaluable. Just as a footballer bounces back from a crushing defeat, Steve learned to recover from failed deployments or security breaches. Strategic planning, akin to reading an opponent’s tactics, mirrored the importance of roadmaps and architectural design in tech. Leadership, too, played a critical role—captains rallying teams on the field were akin to tech leads guiding projects to success. These transferable skills, Steve noted, were his “secret weapon” in navigating the tech industry’s challenges.
Kubernetes as the Ultimate Orchestrator
The discovery of Kubernetes marked a turning point in Steve’s career. He likened it to orchestrating a world-class football squad, where containers are players, roles are defined by role-based access control, and tactics mirror efficient orchestration. Kubernetes, to Steve, was more than a tool—it was a framework for empowering developers through self-service platforms. By providing guardrails, these platforms enable rapid decision-making without compromising stability, much like a coach guiding a team within a strategic framework.
Steve’s work at Jetstack focuses on building such platforms, drawing on the agility and collaboration learned from football. He highlighted the importance of empowering developers to move quickly within structured environments, avoiding the manual processes that bog down innovation. His role as a trainer, having empowered 8,000 individuals worldwide, reflects his commitment to fostering the next generation of cloud native engineers, a mission rooted in the collaborative spirit of his football days.
Overcoming Setbacks as Opportunities
Central to Steve’s narrative was the idea that setbacks are opportunities in disguise. His injury, though catastrophic, forced a reevaluation of his path, leading to a fulfilling career in tech. He encouraged the audience to view challenges as catalysts for growth, urging them to embrace change and maintain a mindset of continuous learning. In the fast-evolving cloud native landscape, where new projects emerge daily, Steve advised focusing on what truly matters rather than attempting to master the entire ecosystem.
He also emphasized the human element in technology. Collaboration, trust, and communication are as vital in tech as they were on the field. Steve’s creation of the Cloud Native Club, a community to support aspiring engineers, underscores his belief that success is amplified by community. By sharing his journey, he invited others to see their challenges as stepping stones, reinforcing that the journey is as significant as the destination.