Posts Tagged ‘CloudNative’
[KCDUK2024] CVEs and Kubernetes: A Love Story? | Marcus Tenorio
In a lively lightning talk at KCDUK2024, Marcus Tenorio, an engineering manager with a background in incident response, brought a fresh perspective on the relationship between Kubernetes and Common Vulnerabilities and Exposures (CVEs). With a nod to the conference’s community spirit, Marcus framed security challenges as opportunities for growth, likening the evolution of Kubernetes security to a love story where vulnerabilities drive collaboration and improvement.
The Evolution of Kubernetes Security
Marcus began by exploring the CVE landscape, drawing from the official CVE feed, MITRE, and NVD databases. He noted that while Kubernetes, launched in 2014, has seen a rise in reported CVEs, this reflects increased scrutiny rather than declining security. Early vulnerabilities, often identified by community members like a Google engineer on GitHub, showcased the power of open-source collaboration. Marcus highlighted that critical CVEs in Kubernetes are relatively rare, contrasting with infamous incidents like Log4j, suggesting a stable core.
He analyzed a sample of 55 CVEs, revealing that the growth in reported vulnerabilities corresponds to Kubernetes’ maturity. As the platform evolves, the community actively identifies and resolves issues, strengthening its security posture. Marcus emphasized that this process mirrors a relationship where challenges foster growth, with each CVE contributing to a more robust ecosystem.
Community-Driven Security
The heart of Marcus’s talk was the role of community in Kubernetes security. He shared an anecdote from an e-commerce platform where a team’s proactive vulnerability hunting led to safer systems, not because vulnerabilities were abundant, but because they were addressed collaboratively. This approach, rooted in policies and shared learning, transforms potential threats into opportunities for improvement.
Marcus encouraged attendees to embrace this “love” for security by fostering open communication and leveraging data to understand vulnerabilities. Tools like Datadog, despite occasional AI hallucinations, help teams analyze and respond to CVEs effectively. By viewing security as a collective journey, Marcus underscored how Kubernetes’ community-driven model drives resilience, aligning with KCDUK2024’s ethos of collaboration.
[KCDUK2024] Platform Orchestrators: The Missing Middle of Internal Developer Platforms | Daniel Bryant
At KCDUK2024, Daniel Bryant, a seasoned platform engineering advocate, delivered a compelling case for platform orchestrators as the critical “missing middle” in internal developer platforms (IDPs). Drawing from his extensive experience with Kubernetes, Mesos, and tools like Backstage and Crossplane, Daniel explored how orchestrators bridge the gap between developer-facing portals and infrastructure layers, enabling scalable, efficient, and user-centric platforms. His talk offered a blueprint for organizations to balance speed, safety, and scale in their platform engineering efforts.
The Evolution of Platform Engineering
Daniel began by contrasting three approaches to platform building: top-down, app-centric portals; bottom-up, infrastructure-focused solutions; and a middle-out, platform-engineering-focused model. Top-down approaches, like Backstage, excel at providing quick wins with developer portals but struggle with day-two operations like upgrades and maintenance. Bottom-up approaches, such as Terraform or Crossplane, offer robust automation but often overwhelm developers with infrastructure complexity. The middle-out approach, which Daniel champions, treats platforms as products, prioritizing user needs and process automation.
He referenced Gartner’s platform engineering model, which identifies three layers: application choreography, platform orchestration, and infrastructure composition. The platform orchestration layer, often overlooked, manages the platform’s lifecycle and APIs, ensuring seamless integration between developer workflows and infrastructure. Daniel’s experience with tools like Crossplane and CNOE (Cloud Native Operational Excellence) highlighted how orchestrators codify business processes, reducing coordination overhead and enabling scalability.
Addressing Developer Pain Points
Modern software engineering faces challenges like slow delivery, high-risk deployments, and tech sprawl. Daniel cited statistics showing that 50% of organizations deploy code less than once a month, and 42% of developers fear production failures. Platform orchestrators address these by offering “everything as a service,” from databases to domain-specific services like fraud detection in finance. By automating manual processes, such as security sign-offs, orchestrators enhance safety and efficiency, allowing developers to focus on coding.
Daniel emphasized the importance of progressive disclosure—presenting simple interfaces initially while enabling advanced functionality as needed. He recounted a past experience where a 500-line YAML configuration overwhelmed developers, underscoring the need for intuitive abstractions. Tools like Open Application Model (OAM) and Score, donated to the CNCF, provide developer-friendly APIs, while orchestrators like Kritik and Cusion Stack manage complex workflows, ensuring platforms remain adaptable to changing business needs.
Building Platforms as Products
The heart of Daniel’s message was treating platforms as products, designed with user needs at the forefront. He advocated for clear domain boundaries, inspired by principles like SOLID and CUPID, to ensure platforms are composable and maintainable. Tools like Kritik, where Daniel contributes, use Kubernetes CRDs to define platform APIs, allowing workflows to be containerized and reusable. This approach enables teams to manage platform components at scale, from rolling out security fixes to integrating auditing processes.
Drawing from Team Topologies, Daniel stressed collaboration between platform and development teams to align on goals like adoption rates and onboarding times. He warned against the “build it and they will come” mentality, urging platform engineers to engage users early and measure success through leading indicators like onboarding efficiency and lagging indicators like incident reduction. By treating platforms as products, organizations can achieve the speed, safety, and scale needed to thrive in cloud-native environments.
Links:
[KCDUK2024] The Joy of DevEx: Tightening Developer Feedback Loops in Kubernetes | Matthew Revell-Gordon
At KCDUK2024, Matthew Revell-Gordon, a senior platform engineering consultant at OpenCredo, shared an engaging exploration of how to enhance the developer experience (DevEx) by tightening feedback loops in Kubernetes. His talk addressed the challenges of local testing for cloud-native applications, offering practical solutions to align local Kubernetes environments with cloud deployments. By leveraging tools like Colima, KIND, and MetalLB, Matthew demonstrated how to create a robust local setup that mirrors production, fostering faster development cycles and a more joyful DevEx.
Challenges of Local Testing in Cloud-Native Environments
Cloud-native technologies like Kubernetes are designed to thrive in distributed, scalable cloud environments, which poses a significant hurdle for local testing. Matthew highlighted that while unit tests and linting are straightforward on a laptop, integration testing and infrastructure configuration validation are far more complex. Laptops lack the scale and fidelity of cloud environments, leading developers to rely heavily on cloud-based testing. This approach, however, introduces lengthy feedback loops, as CI/CD pipelines can take time to queue, bootstrap, and execute—often failing due to trivial errors like typos that could have been caught earlier.
Matthew recounted a real-world scenario where developers used Docker Compose for local testing, only to encounter issues when deploying to Kubernetes due to untested Ingress controller changes. This mismatch delayed releases and frustrated teams, underscoring the need for a local environment that closely emulates production. He emphasized that while cloud testing is inevitable, optimizing local setups can significantly reduce cycle times, catching errors before they reach costly cloud pipelines.
Crafting a Production-Like Local Kubernetes Setup
To address these challenges, Matthew proposed a solution centered on Colima, KIND, and MetalLB to create a local Kubernetes cluster that mirrors cloud deployments. Colima, a lightweight alternative to Docker Desktop, provides a configurable Linux-based virtual machine, allowing developers to fine-tune their environment. KIND (Kubernetes IN Docker) enables multi-node, high-availability clusters with precise control over Kubernetes versions, ensuring consistency with production. MetalLB replaces cloud-specific load balancer controllers, provisioning external IP addresses for local services.
In a live demo, Matthew provisioned a multi-node Kubernetes cluster using KIND, deploying a simple application with node affinity and a load balancer service accessible directly from his Mac. He addressed networking challenges by configuring routes and IP tables, ensuring seamless connectivity between the host, Colima VM, and KIND network. This setup, while not identical to a cloud environment, offers a close approximation, enabling developers to test complex configurations like node affinity and rollouts locally. Matthew also recommended automating these setups with scripts, making them accessible to developers who may not be infrastructure experts, thus enhancing team efficiency.
Enhancing DevEx Through Automation and Community
The true joy of DevEx, Matthew argued, lies in empowering developers to focus on coding rather than wrestling with infrastructure. By packaging local Kubernetes setups into one-click scripts, platform engineers can democratize access to production-like environments, reducing friction and boosting productivity. He cautioned against anti-patterns like creating local-only manifests or altering production setups for testing, as these diverge from the goal of fidelity. Instead, he advocated an iterative approach, starting with the most pressing pain points and gradually refining the local environment.
Matthew’s passion for community resonated throughout his talk. He encouraged platform engineers to share their tooling and scripts, fostering collaboration and easing the burden on developers. By tightening feedback loops, these solutions not only save time and costs but also cultivate a sense of joy in the development process, aligning with the broader ethos of KCDUK2024’s focus on community-driven innovation.
Links:
[DotJs2024] The Future of Serverless is WebAssembly
Envision a computing paradigm where applications ignite with the swiftness of a spark, unbound by the sluggish boot times of traditional servers, and orchestrated in a symphony of polyglot harmony. David Flanagan, a seasoned software engineer and educator with a storied tenure at Fermyon Technologies, unveiled this vision at dotJS 2024, championing WebAssembly (Wasm) as the linchpin for next-generation serverless architectures. Drawing from his deep immersion in cloud-native ecosystems—from Kubernetes orchestration to edge computing—Flanagan demystified how Wasm’s component model, fortified by WASI Preview 2, ushers in nanosecond-scale invocations, seamless interoperability, and unprecedented portability. This isn’t mere theory; it’s a blueprint for crafting resilient microservices that scale effortlessly across diverse runtimes.
Flanagan’s discourse pivoted on relatability, eschewing abstract metrics for visceral analogies. To grasp nanoseconds’ potency—where a single tick equates to a second in a thrashing Metallica riff like “Master of Puppets”—he likened Wasm’s cold-start latency to everyday marvels. Traditional JavaScript functions, mired in milliseconds, mirror a leisurely coffee brew at Starbucks; Wasm, conversely, evokes an espresso shot from a high-end machine, frothy and instantaneous. Benchmarks underscore this: Spin, Fermyon’s runtime, clocks in at 200 nanoseconds versus AWS Lambda’s 100-500 milliseconds, a gulf vast enough to render prior serverless pains obsolete. Yet, beyond velocity lies versatility—Wasm’s binary format, agnostic to origin languages, enables Rust, Go, Zig, or TypeScript modules to converse fluidly via standardized interfaces, dismantling silos that once plagued polyglot teams.
At the core lies the WIT (WebAssembly Interface Types) component model, a contractual scaffold ensuring type-safe handoffs. Flanagan illustrated with a Spin-powered API: a Rust greeter module yields to a TypeScript processor, each oblivious to the other’s internals yet synchronized via WIT-defined payloads. This modularity extends to stateful persistence—key-value stores mirroring Redis or SQLite datastores—without tethering to vendor lock-in. Cron scheduling, WebSocket subscriptions, even LLM inferences via Hugging Face models, integrate natively; a mere TOML tweak provisions MQTT feeds or GPU-accelerated prompts, all sandboxed for ironclad isolation. Flanagan’s live sketches revealed Spin’s developer bliss: CLI scaffolds in seconds, hot-reloading for iterative bliss, and Fermyon Cloud’s edge deployment scaling to zero cost.
This tapestry of traits—rapidity, portability, composability—positions Wasm as serverless’s salvation. Flanagan evoked Drupal’s Wasm incarnation: a full CMS, sans server, piping content through browser-native execution. For edge warriors, it’s liberation; for monoliths, a migration path sans rewrite. As toolchains mature—Wazero for Go, Wasmer for universal hosting—the ecosystem beckons builders to reimagine distributed systems, where functions aren’t fleeting but foundational.
Nanosecond Precision in Practice
Flanagan anchored abstractions in benchmarks, equating Wasm’s 200ns starts to life’s micro-moments—a blink’s brevity amplified across billions of requests. Spin’s plumbing abstracts complexities: TOML configs summon Redis proxies or SQLite veins, yielding KV/SQL APIs that ORMs like Drizzle embrace. This precision cascades to AI: one-liner prompts leverage remote GPUs, democratizing inference without infrastructural toil.
Polyglot Harmony and Extensibility
WIT’s rigor ensures Rust’s safety meshes with Go’s concurrency, TypeScript’s ergonomics—all via declarative interfaces. Spin’s extensibility invites custom components; 200 Rust lines birth integrations, from Wy modules to templated hooks. Flanagan heralded Fermyon Cloud’s provisioning: edge-global, zero-scale, GPU-ready— a canvas for audacious architectures where Wasm weaves the warp and weft.
Links:
[KCDUK2024] Comprehensible Kubernetes: Empowering Scientists with Scalable and Secure Platforms for HPC and AI
At KCDUK2024, Scott Coulton and Tyler, representing StackHPC, delivered an insightful presentation on making Kubernetes accessible to scientists and researchers with minimal technical backgrounds. Their talk focused on crafting cloud-native platforms that prioritize usability, security, and scalability, enabling researchers to focus on their work rather than grappling with complex infrastructure. By leveraging open-source tools like ClusterAPI, Helm, Zenith, and Keycloak, StackHPC has bridged the gap between high-performance computing (HPC) and cloud-native technologies, transforming how scientific research is conducted.
Bridging HPC and Cloud-Native for Research
StackHPC, a Bristol-based consultancy specializing in HPC and cloud solutions, was founded on traditional HPC expertise, such as Slurm clusters and high-performance networking. Scott and Tyler outlined their mission to translate this expertise into the cloud-native realm, primarily for universities and research institutions. Their approach centers on three pillars: reconfigurable infrastructure, performance optimization, and self-service applications. By deploying private OpenStack clouds and Kubernetes clusters, they enable researchers to access tailored environments without needing deep technical knowledge.
The diversity of scientific use cases presents unique challenges. Researchers may require Slurm clusters for batch processing, JupyterHub for data analysis, or GPU-enabled environments for machine learning. Scott highlighted a common thread: scientists are not platform engineers and seek intuitive, reliable platforms. StackHPC’s solution, the LOKI stack (Linux, OpenStack, Kubernetes Infrastructure), integrates open-source tools to provide a flexible, self-service platform that meets these needs while maintaining security and scalability.
The LOKI Stack: A Scalable Solution
Tyler delved into the technical underpinnings of StackHPC’s LOKI stack, which combines ClusterAPI, Zenith, and Azimuth to deliver seamless infrastructure management. ClusterAPI, a declarative API, simplifies Kubernetes cluster lifecycle management, supporting auto-healing and auto-scaling across providers like OpenStack. Tyler explained how Helm charts streamline cluster provisioning, ensuring consistency and ease of deployment. This approach allows operators to manage infrastructure efficiently, freeing researchers from administrative burdens.
Zenith, an innovative application proxy, enables secure external access to applications without public IPs, using SSH tunneling and OIDC authentication. Azimuth, a self-service web portal, offers pre-configured appliances like JupyterHub and Slurm clusters, customizable via Helm charts. Keycloak integration ensures secure access management, allowing platform administrators to create user accounts without granting cloud access. This architecture empowers researchers to deploy and manage platforms independently, aligning complexity with their expertise.
Case Studies: Real-World Impact
Scott presented two case studies illustrating the LOKI stack’s versatility. The first involved deploying JupyterHub for training courses on a private OpenStack cloud. Using Azimuth, administrators created isolated Keycloak realms to manage attendee accounts, ensuring secure access via Zenith’s proxying. This setup allowed course tutors to focus on teaching, with attendees accessing familiar JupyterHub environments without needing cloud credentials. The solution’s simplicity and security made it ideal for educational settings.
The second case study addressed a university’s request for a privacy-preserving large language model (LLM) service. StackHPC developed a Helm chart to deploy open-source LLMs, integrated with Azimuth for ad-hoc testing and ArgoCD for production-grade management. The service, designed to be GDPR-compliant, provided an anonymous interface for staff and students to experiment with LLMs. Monitoring via Prometheus and Grafana ensured reliability, demonstrating how StackHPC’s stack adapts to emerging technologies while maintaining robustness.
Empowering Researchers Through Simplicity
The core takeaway from Scott and Tyler’s talk was that scientists prioritize research over infrastructure management. By offering a platform that balances infrastructure-as-a-service flexibility with platform-as-a-service simplicity, StackHPC empowers researchers to work efficiently. Their commitment to open-source, evidenced by plans to donate Zenith and Azimuth to the CNCF sandbox, underscores their dedication to community-driven innovation. The LOKI stack’s ability to abstract complexity while preserving functionality positions it as a transformative tool for scientific computing.
Links:
[KCDUK2024] From Free Kicks to Git Commits: Steve Wade’s Journey at KCDUK2024
At KCDUK2024, Steve Wade, a cloud native consultant and trainer at Jetstack, captivated the audience with a narrative of transformation, tracing his remarkable journey from professional football to a thriving career in technology. His talk, a blend of personal reflection and technical insight, illuminated the parallels between orchestrating a football team and building self-service platforms on Kubernetes. Steve’s story is one of resilience, adaptability, and the power of transferable skills, offering a compelling blueprint for navigating career pivots and leveraging past experiences to excel in the tech landscape.
From the Pitch to the Keyboard: A Career Pivot
Steve’s journey began at the tender age of four, scouted for his football potential, a moment that ignited a lifelong passion. His early career was defined by dreams of playing in iconic stadiums like Wembley or in a Champions League final. However, a devastating knee injury at 18—shattering his ACL and ligaments—halted his aspirations, plunging him into an identity crisis. Unable to walk for 18 months, Steve faced a profound setback, spiraling into depression and grappling with the question, “Who am I if not a footballer?” This period of introspection, however, became a catalyst for reinvention.
Inspired by his father’s work in IT, Steve’s curiosity was piqued. From his bed, he began exploring technology, diving into coding and the intricacies of programming languages. This marked the beginning of a significant pivot, transitioning from the physical demands of the football field to the intellectual challenges of the digital realm. His relentless drive, honed on the pitch, translated into a disciplined approach to learning, setting the foundation for his eventual expertise in cloud native technologies.
Lessons from Football: Teamwork, Empathy, and Resilience
Steve eloquently drew parallels between the skills cultivated in football and those essential in technology. On the pitch, teamwork was paramount—11 players working cohesively toward a single goal, with intricate passes leading to success. Similarly, in tech, delivering a product requires collaboration across diverse teams, where no one “scores” alone. Steve emphasized empathy as a cornerstone, recounting how supporting teammates in football translated to understanding developers’ needs in tech. Knowing when a colleague needs support, he argued, transforms good teams into great ones, a principle that resonates in both domains.
Resilience, another lesson from his sporting days, proved invaluable. Just as a footballer bounces back from a crushing defeat, Steve learned to recover from failed deployments or security breaches. Strategic planning, akin to reading an opponent’s tactics, mirrored the importance of roadmaps and architectural design in tech. Leadership, too, played a critical role—captains rallying teams on the field were akin to tech leads guiding projects to success. These transferable skills, Steve noted, were his “secret weapon” in navigating the tech industry’s challenges.
Kubernetes as the Ultimate Orchestrator
The discovery of Kubernetes marked a turning point in Steve’s career. He likened it to orchestrating a world-class football squad, where containers are players, roles are defined by role-based access control, and tactics mirror efficient orchestration. Kubernetes, to Steve, was more than a tool—it was a framework for empowering developers through self-service platforms. By providing guardrails, these platforms enable rapid decision-making without compromising stability, much like a coach guiding a team within a strategic framework.
Steve’s work at Jetstack focuses on building such platforms, drawing on the agility and collaboration learned from football. He highlighted the importance of empowering developers to move quickly within structured environments, avoiding the manual processes that bog down innovation. His role as a trainer, having empowered 8,000 individuals worldwide, reflects his commitment to fostering the next generation of cloud native engineers, a mission rooted in the collaborative spirit of his football days.
Overcoming Setbacks as Opportunities
Central to Steve’s narrative was the idea that setbacks are opportunities in disguise. His injury, though catastrophic, forced a reevaluation of his path, leading to a fulfilling career in tech. He encouraged the audience to view challenges as catalysts for growth, urging them to embrace change and maintain a mindset of continuous learning. In the fast-evolving cloud native landscape, where new projects emerge daily, Steve advised focusing on what truly matters rather than attempting to master the entire ecosystem.
He also emphasized the human element in technology. Collaboration, trust, and communication are as vital in tech as they were on the field. Steve’s creation of the Cloud Native Club, a community to support aspiring engineers, underscores his belief that success is amplified by community. By sharing his journey, he invited others to see their challenges as stepping stones, reinforcing that the journey is as significant as the destination.
Links:
[KCDUK2024] Building the Future, Together | KCDUK2024
From Neopets to Google: A Tech Journey
Cheryl Hung, a Senior Director at Arm, delivered an inspiring keynote at KCDUK2024 titled “Building the Future, Together.” Reflecting on her journey as a cloud native pioneer, Cheryl shared personal anecdotes and community insights, weaving a narrative of resilience and collaboration. Born to a non-technical family of Hong Kong immigrants, Cheryl’s entry into tech was sparked by fascination with Google’s innovative culture and early experiments with Neopets’ HTML coding.
At Cambridge, Cheryl pursued a computer science degree, driven by her ambition to join Google. Despite a failed interview at 19, she persevered, securing a role as a software engineer working on Google Maps’ C++ backend. However, burnout and a lack of fulfillment led her to quit after five years, questioning her place in tech. Cheryl’s rediscovery of passion through community meetups, particularly around Docker and containers, marked a turning point, aligning with her prior exposure to Google’s Borg infrastructure.
Embracing Developer Advocacy
Defying warnings about career risks, Cheryl embraced developer advocacy, joining a startup as employee number 12. There, she honed her public speaking skills and founded the Cloud Native London Meetup, building strong community ties. Her role as a CNCF ambassador, secured through a simple photo submission, amplified her influence, leading to her appointment as Director of Ecosystem at CNCF. Cheryl’s tenure involved extensive travel, delivering keynotes and engaging with cloud native “celebrities” like Kelsey Hightower and Liz Rice, though the 2020 pandemic disrupted this dynamic lifestyle.
Cheryl’s interactions revealed the human side of tech luminaries, from Dan Kohn’s supportive mentorship to Kelsey’s off-stage humor. These experiences underscored the importance of authentic connections in driving innovation. Her narrative highlighted the CNCF’s growth from a nascent organization to a powerhouse managing 190 projects, reflecting the community’s role in shaping cloud native technologies.
Navigating Career Transitions
Pregnancy prompted Cheryl to reassess her role at CNCF, leading to a move to Apple, followed by a swift transition to Arm as Senior Director of Ecosystem. Despite personal challenges, including back-to-back maternity leaves, Cheryl’s commitment to community remained steadfast. Her narrative emphasized the power of relationships in career navigation, as opportunities arose through trusted connections rather than formal applications.
Cheryl’s candid reflection on overcoming burnout and imposter syndrome resonated deeply, illustrating the resilience required in tech. Her move to Arm, driven by an irresistible role, underscored her belief in seizing opportunities that align with personal and professional growth. This journey of adaptation and reinvention highlighted the importance of flexibility in a rapidly evolving industry.
Building a Collaborative Future
Cheryl concluded with a call to action, urging attendees to foster community through connection. Acknowledging her struggle with facial blindness, she emphasized the courage required to engage with strangers, advocating for inclusive spaces where diverse ideas thrive. Her vision for the future centers on collective effort, encouraging practitioners to build together, leveraging events like KCDUK2024 to forge lasting relationships.
Cheryl’s narrative serves as a beacon for aspiring technologists, demonstrating that the future of cloud native lies in community-driven innovation. By sharing her vulnerabilities and triumphs, she inspired attendees to embrace collaboration, ensuring the ecosystem continues to evolve through shared knowledge and diverse perspectives.
Links:
[DevoxxGR2024] The Art of Debugging Inside K8s Environment at Devoxx Greece 2024 by Andrii Soldatenko
At Devoxx Greece 2024, Andrii Soldatenko, a seasoned software engineer and tech evangelist at Dynatrace, delivered an engaging presentation on mastering the art of debugging within Kubernetes (K8s) environments. With a blend of humor, practical insights, and real-world strategies, Andrii illuminated the complexities of troubleshooting cloud-native applications. Drawing from his extensive experience, he provided actionable techniques to enhance debugging efficiency, making the session a valuable resource for developers navigating the intricacies of Kubernetes. His talk emphasized proactive design, robust tooling, and a systematic approach to resolving issues in distributed systems.
The Challenges of Debugging in Kubernetes
Andrii began by acknowledging the inherent difficulties of debugging in modern cloud-native environments. Unlike traditional development, where a local debugger suffices, Kubernetes introduces layers of complexity with containers, pods, and distributed architectures. He humorously outlined his “eight stages of debugging,” from denial (“this can’t happen”) to self-realization (“I wrote this code”), resonating with developers who face similar emotional journeys. These stages underscore the psychological and technical hurdles of troubleshooting in K8s, where issues often stem from accidental complexities like misconfigured resources or network policies.
The dynamic nature of Kubernetes, with its orchestration of pods, nodes, and services, demands a shift in debugging mindset. Andrii emphasized that while writing YAML manifests for K8s is straightforward, ensuring they function as intended is not. He highlighted the absence of comprehensive debugging guides, noting that most literature focuses on deployment rather than troubleshooting. This gap inspired his talk, which aimed to equip developers with practical strategies to diagnose and resolve issues effectively.
Strategies for Effective Debugging
To tackle Kubernetes debugging, Andrii proposed a structured approach, starting with a high-level mind map for assessing pod states. For instance, a pod in a “Pending” state might indicate resource shortages or port conflicts, while a “Crashing” pod could signal health probe failures. He focused on scenarios where pods are running but behaving unexpectedly, a common yet challenging issue. Andrii advocated revisiting init containers, which perform setup tasks like data migrations. By temporarily replacing their commands with a sleep directive, developers can use kubectl exec to inspect the container’s state, checking volumes, permissions, or network access.
For containers lacking debugging tools, Andrii introduced ephemeral containers, a Kubernetes feature since version 1.8 designed for interactive troubleshooting. By launching an ephemeral container with tools like netcat or a debugger, developers can inspect a pod’s state without altering its primary container. He shared a practical example of debugging a Go application by sharing process namespaces, allowing access to the application’s processes. This approach enables setting breakpoints and navigating code, even in minimal, distroless containers.
Leveraging Tools for Enhanced Debugging
Andrii showcased several tools to streamline Kubernetes debugging. He recommended building custom debug containers tailored to specific needs, such as including sqlite, python, or network utilities, and shared his own debug container on GitHub. For network-related issues, he highlighted a pre-existing container with tools like tcpdump, which simplifies packet inspection without requiring manual installations. Andrii also praised Stern, a CLI tool for tailing logs across multiple pods in a replica set, making it easier to trace requests and identify exceptions.
For developers using Visual Studio Code, Andrii demonstrated remote debugging by configuring a launch.json file to connect to a Kubernetes pod. By exposing a debug port and using tools like Telepresence, developers can intercept cluster traffic and test changes locally, bypassing slow CI/CD cycles. He also highlighted K9s, a terminal-based UI for Kubernetes, with a custom plugin for initiating debug sessions via kubectl debug. These tools collectively enhance efficiency, allowing developers to focus on problem-solving rather than manual configuration.
Best Practices for Proactive Debugging
Andrii concluded with actionable best practices to prevent and address debugging challenges. He stressed embedding version information, like Git commit SHAs, into container images to synchronize codebases during remote debugging. Scaling down traffic to a single pod ensures consistent debugging sessions, avoiding request distribution across replicas. He also advocated for a blameless culture, where developers use debuggers to slow down and analyze issues methodically rather than rushing to fix symptoms.
By sharing his GitHub repository and additional resources, Andrii encouraged attendees to experiment with these techniques. His talk was a compelling call to action for developers to embrace robust debugging practices, ensuring resilience and reliability in Kubernetes environments. Through practical demonstrations and a lighthearted approach, he demystified the complexities of cloud-native debugging, empowering developers to tackle issues with confidence.
Links:
[KCDUK2024] Sustainability Chronicles: Innovate Through Green Technology With Kepler and KEDA | KCDUK2024
The Imperative of Cloud Sustainability
At KCDUK2024, Katie Gamanji, a Senior Field Engineer at Apple and a prominent figure in the Cloud Native Computing Foundation (CNCF), delivered a compelling session titled “Sustainability Chronicles: Innovate Through Green Technology With Kepler and KEDA.” Katie’s presentation addressed the critical need for environmental consciousness in the tech sector, particularly within the cloud native ecosystem. With a decade of Kubernetes driving industry transformation, Katie emphasized the urgency of integrating sustainability into technological decision-making, given the tech sector’s 1.4% contribution to global greenhouse gas emissions.
Katie outlined the global context, referencing the Paris Agreement (COP21) and the United Nations’ Sustainable Development Goals (SDG13), which call for proactive climate action. She highlighted that adopting renewable energy could reduce tech emissions by 80%, a potential that major cloud providers like GCP, AWS, and Azure are pursuing through net-zero targets. Katie’s narrative framed sustainability as an integral part of operational efficiency, introducing the concept of GreenOps, which aligns resource optimization with reduced environmental impact.
Measuring Emissions with Kepler
Katie introduced Kepler, a CNCF sandbox project developed by Red Hat and IBM, as a pivotal tool for measuring application emissions in Kubernetes. Kepler leverages eBPF to collect energy consumption data, converting it into CO2 emissions using hardcoded emission factors for coal, petroleum, and natural gas. Deployed on a local Kind cluster, Kepler’s Grafana dashboards provide granular insights into emissions per container, day, and namespace. This capability enables organizations to identify high-emission workloads and implement targeted optimizations.
Katie encouraged community engagement with Kepler, noting its sandbox status and potential for growth. By providing feedback or contributing features, practitioners can enhance Kepler’s utility, fostering a collaborative approach to sustainability. Her demonstration underscored the importance of observability in addressing emissions, setting the stage for proactive workload management.
Scaling with KEDA and Carbon-Aware Operator
To move beyond measurement, Katie introduced KEDA (Kubernetes Event-Driven Autoscaler) and its Carbon-Aware Operator, which optimize application scaling based on carbon intensity. KEDA, a graduated CNCF project, scales workloads in response to external events, such as carbon intensity data from grid providers like WattTime or Electricity Map. The Carbon-Aware Operator adjusts replica counts inversely to carbon intensity, maximizing replicas when renewable energy is abundant and minimizing them during high-emission periods.
Katie’s local deployment showcased this dynamic scaling, visualized through Grafana, where low carbon intensity correlated with higher replica counts. This approach exemplifies proactive sustainability, aligning computational demands with environmental impact. Katie’s advocacy for KEDA highlighted its evolution from a niche solution to an industry standard, encouraging practitioners to explore its potential for green innovation.
Community-Driven Sustainability
Beyond technology, Katie emphasized the role of community in driving sustainability. The Kubernetes community’s inclusive ethos fosters innovation through open governance and diverse contributions. She highlighted the CNCF’s Technical Advisory Group (TAG) for Environmental Sustainability, which promotes green practices across projects. Initiatives like the Green Reviews Working Group and sustainability white papers offer avenues for engagement, encouraging practitioners to contribute beyond code.
Katie’s call to action urged attendees to diversify technical boards and create welcoming spaces for new ideas. Her narrative wove together technological innovation and community collaboration, positioning sustainability as a shared responsibility. By integrating tools like Kepler and KEDA with community efforts, the cloud native ecosystem can lead in green technology, shaping a sustainable future.
Links:
[DevoxxGR2024] Devoxx Greece 2024 Sustainability Chronicles: Innovate Through Green Technology With Kepler and KEDA
At Devoxx Greece 2024, Katie Gamanji, a senior field engineer at Apple and a technical oversight committee member for the Cloud Native Computing Foundation (CNCF), delivered a compelling presentation on advancing environmental sustainability within the cloud-native ecosystem. With Kubernetes celebrating its tenth anniversary, Katie emphasized the urgent need for technologists to integrate green practices into their infrastructure strategies. Her talk explored how tools like Kepler and KEDA’s carbon-aware operator enable practitioners to measure and mitigate carbon emissions, while fostering a vibrant, inclusive community to drive these efforts forward. Drawing from her extensive experience and leadership in the CNCF, Katie provided a roadmap for aligning technological innovation with climate responsibility.
The Imperative of Cloud Sustainability
Katie began by underscoring the critical role of sustainability in the tech sector, particularly given the industry’s contribution to global greenhouse gas emissions. She highlighted that the tech sector accounts for 1.4% of global emissions, a figure that could soar to 10% within a decade without intervention. However, by leveraging renewable energy, emissions could be reduced by up to 80%. International agreements like COP21 and the United Nations’ Sustainable Development Goals (SDGs) have spurred national regulations, compelling organizations to assess and report their carbon footprints. Major cloud providers, such as Google Cloud Platform (GCP), have set ambitious net-zero targets, with GCP already operating on renewable energy since 2022. Yet, Katie stressed that sustainability cannot be outsourced solely to cloud providers; organizations must embed these principles internally.
The emergence of “GreenOps,” inspired by FinOps, encapsulates the processes, tools, and cultural shifts needed to achieve digital sustainability. By optimizing infrastructure—through strategies like using spot instances or serverless architectures—organizations can reduce both costs and emissions. Katie introduced a four-phase strategy proposed by the FinOps Foundation’s Environmental Sustainability Working Group: awareness, discovery, roadmap, and execution. This framework encourages organizations to educate stakeholders, benchmark emissions, implement automated tools, and iteratively pursue ambitious sustainability goals.
Measuring Emissions with Kepler
To address emissions within Kubernetes clusters, Katie introduced Kepler, a CNCF sandbox project developed by Red Hat and IBM. Kepler, a Kubernetes Efficient Power Level Exporter, utilizes eBPF to probe system statistics and export power consumption metrics to Prometheus for visualization in tools like Grafana. Deployed as a daemon set, Kepler collects node- and container-level metrics, focusing on power usage and resource utilization. By tracing CPU performance counters and Linux kernel trace points, it calculates energy consumption in joules, converting this to kilowatt-hours and multiplying by region-specific emission factors for gases like coal, petroleum, and natural gas.
Katie demonstrated Kepler’s practical application using a Grafana dashboard, which displayed emissions per gas and allowed granular analysis by container, day, or namespace. This visibility enables organizations to identify high-emission components, such as during traffic spikes, and optimize accordingly. As a sandbox project, Kepler is gaining momentum, and Katie encouraged attendees to explore it, provide feedback, or contribute to its development, reinforcing its potential to establish a baseline for carbon accounting in cloud-native environments.
Scaling Sustainably with KEDA’s Carbon-Aware Operator
Complementing Kepler’s observational capabilities, Katie introduced KEDA (Kubernetes Event-Driven Autoscaler), a graduated CNCF project, and its carbon-aware operator. KEDA, created by Microsoft and Red Hat, scales applications based on external events, offering a rich catalog of triggers. The carbon-aware operator optimizes emissions by scaling applications according to carbon intensity—grams of CO2 equivalent emitted per kilowatt-hour consumed. In scenarios where infrastructure is powered by renewable sources like solar or wind, carbon intensity approaches zero, allowing for maximum application replicas. Conversely, high carbon intensity, such as from coal-based energy, prompts scaling down to minimize emissions.
Katie illustrated this with a custom resource definition (CRD) that configures scaling behavior based on carbon intensity forecasts from providers like WattTime or Electricity Maps. In her demo, a Grafana dashboard showed an application scaling from 15 replicas at a carbon intensity of 530 to a single replica at 580, dynamically responding to grid data. This proactive approach ensures sustainability is embedded in scheduling decisions, aligning resource usage with environmental impact.
Nurturing a Sustainable Community
Beyond technology, Katie emphasized the pivotal role of the Kubernetes community in driving sustainability. Operating on principles of inclusivity, open governance, and transparency, the community fosters innovation through technical advisory groups (TAGs) focused on domains like observability, security, and environmental sustainability. The TAG Environmental Sustainability, established just over a year ago, aims to benchmark emissions across graduated CNCF projects, raising awareness and encouraging greener practices.
To sustain this momentum, Katie highlighted the need for education and upskilling. Resources like the Kubernetes and Cloud Native Associate (KCNA) certification and her own Cloud Native Fundamentals course on Udacity lower entry barriers for newcomers. By diversifying technical and governing boards, the community can continue to evolve, ensuring it scales alongside technological advancements. Katie’s vision is a cloud-native ecosystem where innovation and sustainability coexist, supported by a nurturing, inclusive community.
Conclusion
Katie Gamanji’s presentation at Devoxx Greece 2024 was a clarion call for technologists to prioritize environmental sustainability. By leveraging tools like Kepler and KEDA’s carbon-aware operator, practitioners can measure and mitigate emissions within Kubernetes clusters, aligning infrastructure with climate goals. Equally important is the community’s role in fostering education, inclusivity, and collaboration to sustain these efforts. Katie’s insights, grounded in her leadership at Apple and the CNCF, offer a blueprint for innovating through green technology while building a resilient, forward-thinking ecosystem.