Recent Posts
Archives

Posts Tagged ‘Kubernetes’

PostHeaderIcon [AWSReInvent2025] From Legacy EC2 to Modern EKS: The Tipalti Transformation to Windows Containers

Lecturer

Aiden is a Senior Solutions Architect at AWS, focusing on the modernization of Windows workloads and high-availability container strategies. He has extensive experience helping fintech enterprises transition away from legacy virtualization models. Maya Morv Freeman is an AWS Technical Account Manager who serves as a primary advisor to Tipalti on cloud governance and architectural excellence. Denny Teller is the Lead DevOps Architect at Tipalti, where he oversees the global infrastructure for the company’s payment automation platform. Denny is a pioneer in implementing GitOps and containerization for complex, regulated Windows environments.

Abstract

For many growing enterprises, legacy Windows applications are “constrained” by the scaling limitations and high operational overhead associated with traditional virtual machines. This article examines Tipalti’s successful migration from a monolithic Amazon EC2-based architecture to a highly scalable, containerized solution on Amazon Elastic Kubernetes Service (EKS). The methodology focuses on the implementation of Windows Containers, which allowed Tipalti to achieve a 50% performance improvement while enabling the adoption of advanced auto-scaling and GitOps workflows. The analysis explores the technical challenges of managing process-heavy Windows workloads, the integration of custom logging and monitoring solutions, and the shift toward an immutable infrastructure model. This transformation has provided Tipalti with a resilient foundation for continuous modernization in the competitive fintech market.

The Evolution of Compute: Overcoming the Limitations of Virtualization

The history of enterprise Windows computing has long been defined by an inefficient “one app per server” model, which often led to significant hardware waste and high management costs. While the introduction of hypervisors and virtual machines improved hardware utilization, these systems still carried the heavy overhead of running multiple full operating system instances for every application. Tipalti recognized that to support their rapid global expansion, they needed to move beyond the constraints of traditional Amazon EC2 instances.

The transition to Windows Containers represents the next critical phase in this evolution. Unlike virtual machines, containers share the host’s kernel, which drastically reduces the resource footprint and allows for much higher density on underlying hardware. This efficiency is paired with improved portability, ensuring that the application environment remains identical from a developer’s local machine to the production EKS cluster. For a fintech company like Tipalti, the most vital benefit of this shift is agility; containers can be spun up or down in seconds, allowing the infrastructure to respond instantly to the volatile traffic patterns inherent in global payment processing.

Methodology: Modernizing the Fintech Infrastructure

Tipalti’s transformation followed a rigorous technical roadmap that sought to move their infrastructure from a “legacy” state of manual server management to a “modern” state of automated orchestration. A central component of this strategy was the use of Amazon EKS for Windows, which allowed the team to manage both Linux and Windows workloads through a unified Kubernetes control plane. This eliminated the need for separate management tools and simplified the overall operational landscape.

The implementation methodology addressed several specific Windows-related challenges. Because many of Tipalti’s legacy applications were not originally designed for the ephemeral nature of containers, the team had to implement sophisticated process management techniques. Furthermore, the adoption of GitOps workflows ensured that the entire infrastructure could be managed as code. In this model, every change to the environment is tracked in a version control system and automatically deployed to the cluster, providing a clear audit trail and reducing the risk of human error. To ensure complete visibility, the team also developed custom logging and monitoring solutions tailored to the telemetry requirements of Windows containers, ensuring that the DevOps team could maintain high availability even during rapid deployment cycles.

Technical Analysis of Performance and Scalability Gains

The move to a containerized EKS environment delivered immediate and measurable technical advantages for Tipalti. One of the most significant outcomes was a documented 50% performance improvement for core payment processing services. This gain was achieved through more efficient resource allocation and the ability to leverage Kubernetes’ native auto-scaling capabilities, which ensure that compute power is always perfectly matched to the current workload.

Operational simplicity also improved as the team moved away from the administrative burden of patching and maintaining hundreds of individual EC2 instances. By using container images, Tipalti shifted toward an immutable infrastructure model, where updates are performed by replacing containers rather than modifying them in place. This has resulted in better “bin-packing,” where more applications are packed onto fewer EC2 nodes, leading to substantial cost savings without compromising on throughput or reliability. A technical hurdle overcome during this process involved managing legacy Windows behaviors that expected persistent file systems; this was resolved by integrating modern Container Storage Interface (CSI) drivers that provide persistent storage to ephemeral containers.

Consequences: Establishing a Foundation for Continuous Innovation

For Tipalti, the successful implementation of Windows containers was not viewed as a final destination but rather as the essential foundation for continuous modernization. By adopting Kubernetes, the organization has unlocked several strategic advantages. They are now able to implement the most modern DevOps practices and tools, which are natively designed for containerized ecosystems. This has significantly accelerated their release cycles and improved the overall quality of their software.

Furthermore, the new infrastructure is inherently more resilient. The automated health checks and self-healing properties of Amazon EKS ensure that the global payment system remains available 24/7, even in the event of hardware failure. Most importantly, the platform is now “future-ready.” Having a containerized environment makes it far easier to integrate advanced cloud-native services, such as AI-driven fraud detection or serverless functions, which would have been prohibitively difficult to implement in the previous VM-based architecture. Tipalti’s journey demonstrates that modernizing the compute layer is the primary enabler for broader business innovation.

Conclusion

The journey of Tipalti from Amazon EC2 to Amazon EKS provides a definitive roadmap for any enterprise seeking to modernize legacy Windows applications. By embracing the efficiency of Windows containers and the power of Kubernetes orchestration, Tipalti has transformed a traditionally rigid system into a high-performance engine for global fintech growth. Their experience highlights that successful modernization requires a combination of strategic technical decisions, a commitment to DevOps excellence, and a focus on long-term scalability. This transformation proves that even the most “constrained” legacy applications can be revitalized to meet the demands of the modern digital economy.

Links:

PostHeaderIcon [VoxxedDaysLuxemburg2026] Deploying Often, Stressing Less: Architecting Critical Production Feature Flags

Lecturers

The lecture was co-delivered by Marion Chineaud and Elise Souvannavong, both Full Stack Developers at Takima, a French software engineering and consulting firm. Marion and Elise specialize in building robust, high-volume Java/Spring and Angular applications and implementing modern DevOps practices, including trunk-based development and continuous deployment.

Abstract

In modern software engineering, delaying releases until large feature sets are completed introduces integration risks, complex rebasing conflicts, and production instability. This talk addresses how to exit the binary “all-or-nothing” deployment model by implementing feature toggling and Trunk-Based Development. Grounded in real-world scenarios from Takiship—a logistics microservices ecosystem built with Java Spring, Angular, Kubernetes, and ArgoCD—the session outlines five architectural flag categories: Release Flags, Ops Flags, Experimental Flags (A/B Testing), Shadow Toggles, and Canary Releases.

The speakers detail practical implementation patterns—ranging from Spring properties and @RefreshScope to database-backed administration panels—while confronting the operational overhead of flag pollution. Finally, the presentation connects deployment frequency directly to Google’s DORA metrics, demonstrating how structured flag lifecycles create a robust safety net for modern continuous delivery and AI-assisted development workflows.

The Monolithic Branching Dilemma vs. Trunk-Based Development

Traditional GitFlow strategies often isolate large features on long-lived branches over extended periods. When multiple engineers alter overlapping microservices, merging results in severe rebase friction, missed edge cases, and high-risk releases.

+---------------------------------------------+
|        Traditional GitFlow Risk             |
|                                             |
|  Dev Branch 1: [--- 2 Months Dev ---]       |
|                                     \       |
|  Dev Branch 2: [--- Rebase Friction -\--->  |
|                                       \     |
|  Main Branch:  ========================(FAIL)
+---------------------------------------------+
|        Trunk-Based + Feature Flags          |
|                                             |
|  Small Batch:  --+---+---+---+---> Main     |
|                  |   |   |   |              |
|  Release Flag:  [OFF][OFF][OFF][ON]         |
+---------------------------------------------+

Transitioning to Trunk-Based Development shortens iteration cycles. Code is integrated into the main branch frequently in small batches. To prevent incomplete features from exposing half-finished functionality to end-users, teams decouple physical code deployment from logical feature activation through Release Flags.

Core Benefits of Feature Toggling

  • Decoupled Lifecycle: Code can be safely pushed to production while dormant, awaiting business approval or QA validation.
  • Instant Rollbacks: When an incident occurs in production, disabling a flag replaces frantic hotfix deployments with an instant configuration change.
  • Granular Task Decomposition: Epics spanning multiple microservices can be split into small, trackable tasks that can be developed, merged, and tested in parallel.

Architectural Taxonomy of Feature Flags

Feature flags are not monolithic; they serve distinct technical and business stakeholders across different lifecycles.

Flag Type Primary Target Audience Core Operational Purpose Lifespan Strategy
Release Flag Developers / QA / Product Decouples code deployment from feature activation. Temporary: Removed after feature adoption.
Ops Flag Systems / DevOps Engineers Dynamic runtime throttling, pagination bounds, and kill-switches. Permanent: Retained indefinitely for operational control.
Experimental Flag Product Managers / Data Analysts A/B testing user interface variants and statistical conversion paths. Temporary: Cleaned up after data collection ends.
Shadow Flag Core Engineering Teams Zero-tolerance validation via silent dual-execution on production traffic. Temporary: Removed post-algorithm validation.
Canary Flag Product & Support Teams Gradual percentage rollouts and progressive audience segmentation. Temporary: Decommissioned after 100% rollout.

Technical Implementations in Java Spring & Kubernetes

Depending on security constraints and autonomy requirements, flag state management can be implemented across three distinct layers.

+---------------------------------------------+
|          Flag Management Taxonomy           |
|                                             |
|  1. Static Application Configuration        |
|     - Spring @ConfigurationProperties       |
|                                             |
|  2. Dynamic GitOps & Hot Reloading          |
|     - Kubernetes ConfigMaps + ArgoCD        |
|     - Spring Cloud @RefreshScope Proxy      |
|                                             |
|  3. DB-Backed Admin Portal                  |
|     - Relational Toggles Table              |
|     - REST Control Endpoints (GET/PUT)      |
+---------------------------------------------+

Option 1: Static Application YAML Configuration

The simplest implementation encapsulates toggles into dedicated configuration properties, separating configuration parameters from domain logic.

@Configuration
@ConfigurationProperties(prefix = "delivery.feature")
public class FeatureFlagsProperties {
    private boolean expressDeliveryEnabled;

    public boolean isExpressDeliveryEnabled() {
        return expressDeliveryEnabled;
    }

    public void setExpressDeliveryEnabled(boolean expressDeliveryEnabled) {
        this.expressDeliveryEnabled = expressDeliveryEnabled;
    }
}

@Service
public class ExpressDeliveryService {
    private final FeatureFlagsProperties properties;

    public ExpressDeliveryService(FeatureFlagsProperties properties) {
        this.properties = properties;
    }

    public void processDelivery(Order order) {
        // Centralized evaluation entry point
        if (properties.isExpressDeliveryEnabled()) {
            executeExpressWorkflow(order);
        } else {
            executeStandardWorkflow(order);
        }
    }
}

  • Limitation: Toggling state requires a Git commit, triggering full application rebuilding and pod redeployment.

Option 2: Dynamic GitOps with Spring Cloud @RefreshScope

To achieve zero-downtime hot reloading without restarting JVM instances, configuration properties are stored in a dedicated GitOps repository managed by ArgoCD and mapped into Kubernetes ConfigMaps.

@Component
@RefreshScope
@ConfigurationProperties(prefix = "delivery.ops")
public class OpsFlagsProperties {
    private int maxHistoricalFetchDays = 7;

    public int getMaxHistoricalFetchDays() {
        return maxHistoricalFetchDays;
    }

    public void setMaxHistoricalFetchDays(int maxHistoricalFetchDays) {
        this.maxHistoricalFetchDays = maxHistoricalFetchDays;
    }
}

  • Mechanism: Spring Cloud wraps the bean within a dynamic proxy. Invoking POST /actuator/refresh invalidates the target proxy cache, forcing subsequent calls to pull updated parameters directly from the configuration server.
  • Critical Restrictions: @RefreshScope cannot be used on scheduled tasks (@Scheduled) or stateful bean dependencies, as forced cache eviction can cause runtime crashes.

Option 3: Database-Backed Administration Portal

When non-technical stakeholders (Product Managers, QA leads) require direct runtime control, flags can be persisted in a database and modified via an administrative dashboard.

CREATE TABLE feature_toggles (
    id VARCHAR(64) PRIMARY KEY,
    toggle_code VARCHAR(64) NOT NULL UNIQUE,
    is_enabled BOOLEAN NOT NULL DEFAULT FALSE,
    display_title VARCHAR(128) NOT NULL
);

To prevent performance degradation from frequent database checks, frontend applications should retrieve the complete active toggle state array alongside user authentication payloads during initial load.

Operational Scenarios and Deployment Patterns

Scenario 1: Managing Traffic Volatility with Ops Flags

During high-volume periods (such as Black Friday or Cyber Week), legacy queries or third-party dependencies can experience severe performance degradation. Ops Flags convert rigid values into dynamic system tuners.

+---------------------------------------------+
|         Ops Flag Runtime Throttling         |
|                                             |
|  Normal Operations  ---> Fetch 30 Days Data |
|                                             |
|  Traffic Surge      ---> Adjust Ops Flag    |
|                          (Fetch 7 Days Data)|
|                                             |
|  System Degradation ---> Trigger Kill Switch|
|                          (Bypass Dependency)|
+---------------------------------------------+

Instead of hardcoding limits, parameters such as query batch sizes, external API timeouts, retry limits, and database pagination bounds are evaluated at runtime.

Scenario 2: Statistical Validation via A/B Testing

When evaluating architectural or UI choices (e.g., standard form vs. multi-step wizard), teams can run both variants concurrently in production.

+---------------------------------------------+
|        Deterministic A/B Routing            |
|                                             |
|  Inbound Request                            |
|        |                                    |
|        v                                    |
|  [ API Gateway ]                            |
|        |-- Cookie Present? -> Route to Variant
|        |-- Cookie Missing? -> Hash User ID  |
|                                (Assign A/B) |
|        v                                    |
|  [ Sticky Session Cookie Set ]              |
+---------------------------------------------+

  • Sticky Sessions: Random allocation must be bound deterministically using session cookies or hashed user IDs. Once assigned, a user must consistently see the same variant to prevent confusing user experiences.
  • Analytics Integration: Key performance metrics—conversion rates, system performance, error logs, and navigation paths—must be tagged with the active variant ID.

Scenario 3: Zero-Tolerance Domain Changes with Shadow Mode

For critical subsystems where failure presents direct business risk (e.g., billing engine updates), Shadow Mode (Dry-Run) runs both legacy and new implementations concurrently.

+---------------------------------------------+
|           Shadow Mode Execution             |
|                                             |
|                Inbound Transaction          |
|                         |                   |
|           +-------------+-------------+     |
|           |                           |     |
|           v                           v     |
|     [ Legacy Engine ]           [ New Engine ]|
|           |                           |     |
|           v                           v     |
|    (Return Result)             (Log Analytics)|
|           |                           |     |
|           +-------------+-------------+     |
|                         |                   |
|                         v                   |
|             [ Differential Audit Log ]      |
+---------------------------------------------+

  1. Incoming requests enter the production gateway.
  2. The legacy engine calculates the result and returns it directly to the customer.
  3. The secondary engine executes the transaction asynchronously in isolation.
  4. Output values, execution traces, and performance characteristics are sent to diagnostic logging systems to surface unexpected discrepancies.

Scenario 4: Risk Mitigation via Canary Releases

During major infrastructure overhauls (such as simultaneous ORM migrations, frontend framework updates, and UI redesigns), changes can be rolled out progressively across user tiers.

+---------------------------------------------+
|           Canary Release Phasing            |
|                                             |
|  Phase 1:  [ Beta Testers ]                 |
|            -> Collect feedback & telemetry  |
|                                             |
|  Phase 2:  [ Standard B2B Customers ]       |
|            -> Validate performance load     |
|                                             |
|  Phase 3:  [ High-Value Key Accounts ]      |
|            -> Complete feature cutover      |
+---------------------------------------------+

Technical Debt Management and Lifecycle Cleanup Strategy

Unmanaged feature flags can lead to operational complexity. Accumulated flag combinations increase test surface areas, complicate local debugging, and increase code clutter.

+---------------------------------------------+
|         Flag Lifecycle Governance           |
|                                             |
|  Release/Experimental Toggles               |
|  [ Created ] -> [ Validated ] -> [ REMOVED ]|
|                                             |
|  Operational Toggles                        |
|  [ Created ] -> [ Maintained Long-Term ]    |
+---------------------------------------------+

Protocol for Technical Debt Mitigation

  1. Centralized Evaluation: Restrict flag conditional evaluation (if/else) to a single service layer or facade point. Avoid scattering flags across nested domain methods.
  2. Automated Cleanup Tickets: Whenever a new temporary toggle is created, an associated cleanup task must be filed immediately in the sprint backlog.
  3. Traceable Annotations: Include explicit inline markers (e.g., // TODO: TOGGLE_CLEANUP_KEY) within target code repositories to streamline string-search audits.
  4. Lifecycle Separation: Maintain a clear operational distinction between temporary release toggles (which are decommissioned post-rollout) and permanent operational controls.

Industrial Context: DORA Metrics and AI Integration

Continuous delivery performance relies on key operational metrics evaluated by Google’s DevOps Research and Assessment (DORA) team:

  • Deployment Frequency: How often code is successfully deployed to production.
  • Lead Time for Changes: The duration required for a committed feature to reach production.
  • Change Failure Rate: The percentage of deployments causing production defects.
  • Failed Service Recovery Time (MTTR): The time required to restore service stability following an outage.

By decoupling deployment from feature activation, teams can increase deployment frequency while keeping change failure rates low. Toggles also provide an instant recovery mechanism (reducing MTTR) by converting complex rollback procedures into configuration changes.

+---------------------------------------------+
|         AI-Assisted CI/CD Guardrails        |
|                                             |
|  Autonomous Agent Code Generation           |
|                     |                       |
|                     v                       |
|        [ Feature Flag Enclosure ]           |
|                     |                       |
|                     v                       |
|  [ Automated Pipeline & Observability ]     |
|                     |                       |
|             +-------+-------+               |
|             |               |               |
|      (Stable Stream)  (Anomalies)           |
|             |               |               |
|             v               v               |
|      Keep Feature     Disable Flag          |
+---------------------------------------------+

As autonomous AI agents generate larger portions of application code, feature flags serve as a key runtime safety boundary. Enclosing AI-generated code within dynamic feature toggles provides an immediate circuit breaker to isolate anomalies, lower integration costs, and maintain production stability.

Links

PostHeaderIcon [VoxxedDaysBucharest2026] Optimizing LLM Inference on Kubernetes: Abdel Sghiouar on Practical Techniques for the Rest of Us

Lecturer

Abdel Sghiouar is a Developer Advocate at Google Cloud with deep expertise in cloud-native technologies, Kubernetes orchestration, and AI/ML workload optimization. Drawing from a robust background in infrastructure engineering and open source contributions, Abdel helps organizations design, deploy, and tune complex AI applications for production environments across diverse infrastructures.

Abstract

While major cloud providers and hyperscalers leverage virtually unlimited computational resources, the majority of organizations face significant constraints when operationalizing Large Language Models. Abdel Sghiouar presents a comprehensive set of practical strategies for optimizing LLM inference workloads on Kubernetes. The session systematically addresses container and model optimization techniques, accelerator management, data persistence and storage considerations, networking and intelligent load balancing, and advanced observability practices. Emphasis is placed on open-source tools and architectural patterns that deliver meaningful cost-performance improvements adaptable to on-premises, hybrid, and public cloud deployments.

Understanding LLM Inference Characteristics and Challenges

Large Language Models continue their rapid evolution in both scale and sophistication. Architectural innovations such as mixture-of-experts (MoE) enable dynamic activation of specialized sub-networks, while multi-modal capabilities process diverse inputs including text, images, audio, and video. Expanded context windows support richer interactions but demand substantial memory resources.

Inference execution comprises two primary phases with contrasting characteristics: the prefill stage (encoding input tokens, predominantly compute-bound) and the decode stage (token generation, typically memory-bound). KV (key-value) caching optimizes conversational flows by preserving intermediate states, avoiding redundant prefill computations for subsequent messages.

Deployment topologies vary considerably. Single-host single-accelerator setups predominate for local development and experimentation (e.g., using Ollama). Single-host multi-accelerator configurations require model sharding across GPUs within one machine. Multi-host distributed deployments introduce complex requirements for high-bandwidth, low-latency interconnects to maintain coherent context across nodes. Each topology presents distinct challenges regarding scalability, fault tolerance, and operational complexity.

Container, Model, and Storage Optimizations

Inference serving runtimes and model artifacts generate exceptionally large container images, frequently exceeding several gigabytes prior to incorporating weights. Conventional optimization strategies like multi-stage builds or native compilation (e.g., GraalVM) prove inadequate for these workloads.

Distributed caching solutions such as Spiegel provide cluster-wide image and model artifact caching, substantially reducing repeated pulls from external registries. Kubernetes-native features enabling containers as volumes allow separate packaging of models, which can then be mounted efficiently onto serving runtimes. When combined with caching layers, these approaches dramatically accelerate cold starts.

Quantization techniques offer another lever, reducing numerical precision (e.g., FP16 to INT8 or lower) to decrease memory footprints while preserving sufficient accuracy for many applications. Careful selection of quantization levels based on task sensitivity balances performance and quality.

Accelerator Management and Dynamic Resource Allocation

Kubernetes has supported GPU scheduling through device plugins for several years. However, static device configurations struggle with real-world constraints including accelerator scarcity and heterogeneous hardware fleets.

Dynamic Resource Allocation, matured in recent Kubernetes versions, introduces flexible resource claiming based on abstract characteristics rather than rigid device specifications (e.g., requesting “NVIDIA GPU with minimum 30GB memory and specific core count”). This enables more efficient scheduling across mixed clusters and better utilization rates.

Integration with cluster autoscalers allows on-demand provisioning, addressing both availability gaps and cost optimization by scaling resources precisely to workload demands. Platform operators describe device inventories; application teams specify requirements, with the scheduler performing intelligent matching.

Networking, Load Balancing, and Observability Considerations

LLM traffic profiles differ markedly from conventional web workloads. Requests exhibit high variability in size and computational intensity (simple text queries versus multi-modal inputs), while responses frequently involve streaming token generation. Standard round-robin load balancing produces inefficient distributions, with certain backends becoming overloaded while others remain underutilized.

The Kubernetes Gateway API, augmented with custom endpoint selection logic, supports sophisticated routing decisions based on request attributes extracted from bodies (model identifier, input modality, streaming requirements) combined with real-time backend telemetry. This facilitates intelligent traffic steering, prioritization of business-critical workloads, and maintenance of sticky sessions necessary for coherent streaming interactions.

Comprehensive observability must encompass prefill and decode phase latencies, KV cache hit rates, token generation throughput, GPU utilization, and end-to-end request metrics. Integration with Prometheus, Grafana, and specialized LLM monitoring solutions provides actionable insights for capacity planning and bottleneck identification.

Practical Patterns and the LLM-D Project

The LLM-D initiative, hosted under the Linux Foundation with contributions from Google, IBM, NVIDIA, and additional partners, aggregates architectural patterns, performance benchmarks, and reference implementations for production-grade inference. Key elements include optimized prefill/decode separation, advanced routing logic often leveraging engines like vLLM, and comprehensive guidance for multi-node deployments.

A holistic, layered optimization strategy proves most effective: infrastructure-level improvements (caching, persistent volumes), platform capabilities (dynamic scheduling, intelligent networking), and application-level choices (model quantization, serving engine selection). Organizations without hyperscale resources can still achieve competitive efficiency and scalability through disciplined application of these patterns.

Links:

PostHeaderIcon [DevoxxGR2026] GenAI on Kubernetes: Training, Inference, and Serving in Production Environments

Lecturer
Alessandro Vozza is a seasoned cloud-native advocate and technologist with deep expertise in Kubernetes and AI/ML operations. He contributes actively to open-source communities and focuses on practical, scalable deployments of generative AI workloads. As a speaker and practitioner, Alessandro emphasizes operational excellence, resource efficiency, and the integration of modern AI tools within established cloud-native platforms.

Abstract
In this hands-on tutorial at Devoxx Greece 2026, Alessandro Vozza guides developers through the complete lifecycle of running generative AI workloads on Kubernetes. From distributed training jobs with GPU scheduling to optimized inference and scalable model serving, the session demonstrates how to leverage operators, autoscaling, vector stores, and frameworks like KServe, Ray, vLLM, and Kubeflow. Attendees gain actionable insights into designing efficient GPU clusters, fine-tuning models securely, and deploying production-grade architectures that integrate seamlessly with existing Kubernetes expertise.

The Convergence of Kubernetes and Generative AI

Kubernetes has evolved into the de facto platform for orchestrating complex, resource-intensive workloads, including those powered by generative AI. Vozza begins by contextualizing the challenges: training large models demands massive parallel computation across GPUs, inference requires low-latency serving under variable traffic, and the entire pipeline must remain observable, secure, and cost-effective. Traditional approaches struggle with these demands, but Kubernetes patterns—scheduling, autoscaling, and declarative resource management—provide a robust foundation.

The session highlights how the community has responded with specialized tools. Projects like Kubeflow address the full ML lifecycle, while KServe and vLLM focus on high-performance inference. These build upon core Kubernetes capabilities, allowing teams to treat AI workloads with the same rigor applied to microservices.

Distributed Training and GPU Orchestration

Training generative models is computationally intensive and benefits enormously from Kubernetes’ scheduling strengths. Vozza demonstrates launching distributed training jobs, emphasizing GPU-aware scheduling through device plugins and resource requests. Nodes are labeled with GPU capacity, enabling the scheduler to place pods on suitable hardware.

The tutorial covers hyperparameter tuning with tools like Katib, which automates experimentation across multiple configurations. Fine-tuning involves augmenting base models with domain-specific data, a process that Kubernetes orchestrates reliably through persistent volumes and checkpointing. Attendees learn to monitor training progress using built-in observability and handle failures gracefully with retries and job controllers.

Resource efficiency emerges as a key theme. Techniques such as multi-instance GPU (MIG) partitioning allow a single physical GPU to support multiple smaller workloads, maximizing utilization without over-provisioning expensive hardware.

Inference Serving and Model Deployment

Once trained, models must be served efficiently. Vozza walks through deploying inference endpoints with KServe, which abstracts the complexities of scaling and routing. vLLM serves as the high-throughput inference engine, leveraging continuous batching and paged attention for superior performance.

The architecture supports multi-model serving, where a single deployment handles various models based on request characteristics. Gateway API extensions make the ingress layer LLM-aware, enabling intelligent routing based on factors like key-value cache state or model specialization. This ensures optimal resource allocation and minimal latency.

Autoscaling plays a critical role. Horizontal Pod Autoscaler (HPA) combined with KEDA reacts to custom metrics such as queue depth or tokens processed per second, dynamically adjusting replicas to match demand while controlling costs.

Operational Considerations and Best Practices

Production readiness demands comprehensive observability. Vozza integrates Prometheus exporters and logging to track token throughput, latency, and GPU utilization. Security best practices include least-privilege access for model endpoints and encrypted communication.

The tutorial addresses common pitfalls: managing model registries for versioning, handling cold starts through caching, and ensuring reproducibility across environments. By treating models as first-class Kubernetes citizens, teams achieve consistent deployments from development to production.

Practical Roadmap and Future Directions

Participants receive a working reference setup they can adapt immediately. Vozza encourages starting small—perhaps with a single-model inference service—before scaling to distributed training and multi-model architectures. The session reinforces that Kubernetes knowledge directly transfers to AI operations, lowering the barrier for traditional platform teams.

Looking ahead, evolving features like dynamic resource allocation and improved GPU topology awareness will further streamline GenAI workloads. The message is clear: Kubernetes is not merely compatible with generative AI; it is becoming the preferred operational layer for the entire lifecycle.

Links:

PostHeaderIcon [DevoxxFR2026] Common Expression Language (CEL): A Fast, Portable, and Secure Expression Language for Modern Applications

Lecturer

Alex Snaps is a Tech Lead at Red Hat working on the Quarkus project. He maintains the Rust implementation of CEL and contributes to the broader ecosystem, bringing deep expertise in language runtimes, performance, and secure extensibility.

Abstract

Alex Snaps introduces the Common Expression Language (CEL), a domain-agnostic expression language designed for safe, high-performance evaluation within larger applications. Originating from Google, CEL emphasizes strong typing, sandboxed execution, and extensibility while maintaining portability across implementations in Go, Java, C++, Rust, and others. Through syntax exploration, type checking, cost estimation, and practical integration examples, the talk demonstrates why CEL excels for policy enforcement, validation, filtering, and authorization in cloud-native and API-driven environments.

Origins and Design Philosophy of CEL

CEL emerged from Google’s need for a lightweight, embeddable expression evaluator capable of running safely in performance-critical paths. First released around 2017 (with earlier internal variants), it targets scenarios where user-provided or configuration-driven logic must execute with predictable latency and strict safety guarantees. Unlike general-purpose scripting languages, CEL deliberately restricts Turing-completeness to prevent denial-of-service through infinite loops or excessive computation.

Core tenets include:

  • Strong Static Typing: All expressions are type-checked before evaluation.
  • Predictable Performance: Cost estimation and constant folding occur at check time.
  • Portability: Abstract Syntax Tree (AST) format enables cross-language evaluation.
  • Extensibility: Custom functions, macros, and types can be added per domain.

These properties make CEL ideal for Kubernetes (Custom Resource Definition validation), API gateways, authorization systems, and configuration engines.

Syntax and Core Language Features

CEL syntax resembles a blend of C-style expressions and modern collection comprehensions. Basic operations, conditionals, and field access feel familiar:

  • Arithmetic and comparisons
  • Logical operators
  • Ternary expressions
  • Optional chaining with ? and or
  • Collection operations via macros like exists, all, map

Notable features include:

  • Macros: all(resources, r, r.startsWith('email')) binds variables and applies predicates.
  • Optional Navigation: obj.?field.or(0) safely accesses potentially absent fields.
  • Message Construction: Direct construction of Protocol Buffer messages within expressions.
  • Strict Typing: No implicit coercion; uint(1) == 1 fails type checking.

The language integrates seamlessly with Protocol Buffers, treating well-known types like timestamps and durations as first-class citizens.

The Evaluation Pipeline: Parse, Check, Evaluate

CEL processing follows a clear separation optimized for control-plane versus data-plane workloads:

  1. Parse: Validates syntax and produces an AST. Feature flags can disable risky syntax (e.g., optional navigation).
  2. Check: Performs type resolution, overload selection, constant folding, and cost estimation. This phase catches errors early and enables optimization.
  3. Evaluate: Executes the (potentially optimized) AST against a bound environment in the hot path.

Environments declare variables and functions available to expressions. Cost limits prevent expensive evaluations in production.

Portability shines here: an AST checked in one language can be evaluated in another, facilitating polyglot systems.

Extensibility and Real-World Integration

CEL’s power emerges through domain-specific extensions. Custom functions, member overloads, and macros allow tailoring to specific needs without compromising safety.

In the Quarkus/Gateway API context, CEL evaluates policies attached to Kubernetes resources. Expressions navigate complex object graphs, enforce authorization, and implement fine-grained controls. The Rust implementation (maintained by Snaps) demonstrates low-level integration, including trait-based value handling and flexible indexing.

Examples illustrate adding domain functions like isPrime or complex policy logic matching gateways and routes.

Performance, Security, and Ecosystem Maturity

CEL achieves high performance through ahead-of-time type checking, constant folding, and minimal runtime overhead. Implementations in Go and Java (reference) are mature; Rust and others continue evolving toward full specification compliance.

Security model emphasizes sandboxing: no arbitrary code execution, bounded computation, and explicit environment control. This makes CEL suitable for untrusted user input in API filters, validation rules, and authorization decisions.

The ecosystem includes playgrounds, conformance test suites, and codelabs across languages, lowering the barrier to adoption.

Conclusion

Common Expression Language offers a compelling balance of expressiveness, safety, and speed for embedding dynamic logic in applications. Its strong typing, cost awareness, and extensibility address real challenges in cloud-native policy and configuration management. As organizations seek safer alternatives to full scripting engines, CEL provides a mature, battle-tested solution that continues gaining traction across diverse technology stacks.

Links:

PostHeaderIcon [DevoxxBE2025] Not Just Code: Abusing Claude Code for Non-Coding Tasks

Lecturer

Barry van Someren operates a compact DevOps hosting and consulting enterprise named CoffeeSprout ICT Services. Previously engaged as a dedicated Java programmer, he now oversees Java-based systems and develops in-house solutions. Barry positions himself as an expert in averting common operational pitfalls such as memory exhaustion or storage shortages.

Abstract

This article scrutinizes the unconventional deployment of Claude Code, an AI-driven coding aide, in domains extending far beyond software creation. It probes into Barry’s methodologies for leveraging the tool in operational duties, infrastructure orchestration, and ad hoc automations, grounded in tangible scenarios. The examination encompasses the inception of these applications, practical executions, triumphs alongside mishaps, and ramifications for forthcoming AI-facilitated workflows in DevOps landscapes.

Inception and Justification for Extended Applications

The genesis of employing Claude Code for purposes unrelated to programming emerged from routine engagements with large language models in configuration oversight. Barry initially harnessed these models to craft Ansible playbooks, a YAML-centric framework for delineating system states. Ansible facilitates the depiction of desired configurations, enabling automated enforcement across servers. During one such interaction, the model proposed executing a command to ascertain a file path, sparking the realization that Claude could transcend mere suggestion to active participation in debugging and setup on development platforms.

This pivot stems from the acknowledgment that numerous operational elements mirror code structures. Infrastructure configurations, for instance, can be codified, while fleeting assignments may not warrant full-fledged scripting. Recurring chores often reveal themselves post hoc, prompting Barry to instruct Claude to formulate reusable scripts after task completion. Notably, this approach eschews intricate prompt crafting; initiating a dialogue within Claude’s interface, refining directives iteratively, suffices for efficacious outcomes.

Furthermore, the rationale hinges on friction reduction in learning novel utilities. Barry recounts configuring a rudimentary virtual machine, where Claude undertook preparatory steps, thereby expediting assimilation of unfamiliar technologies. This proves particularly advantageous in conference settings like Devoxx, where novel concepts abound, allowing practitioners to experiment swiftly without exhaustive manual setup.

Claude Code’s allure lies in its subscription framework, mitigating earlier credit-based expenditures that could escalate to substantial sums daily. The advent of affordable plans democratizes access, rendering it viable for exploratory uses. Its acumen in encoding and tool proficiency outpaces contemporaries, although rivals like ChatGPT’s Codex narrow the disparity. Consequently, Barry advocates for its adoption in streamlining DevOps, transforming mundane operations into efficient processes.

Methodological Executions and Illustrative Cases

Barry’s technique involves granting Claude terminal access within controlled environs, such as virtual machines or containers, to execute commands and scripts. This necessitates safeguards: employing disposable instances, restricting privileges via non-root users, and isolating sensitive data. For demonstration, he configures a Spring Pet Clinic application on Ubuntu, commencing with package updates and Java installation.

In one instance, Claude autonomously installs PostgreSQL, initializes a database, and integrates it with the application by modifying configuration files. It generates passwords—albeit simplistic ones—and applies them consistently, showcasing its aptitude for cross-file correlation. Another example entails heap analysis on a Java application; Claude employs jmap to capture heap dumps, analyzes them with jhat, and identifies memory leaks, all while navigating command-line intricacies.

A compliance scenario highlights versatility: adhering to energy conservation regulations, Claude devises scripts to throttle CPU frequencies during off-hours, generates audit logs, and verifies adherence, yielding a 15% reduction in power consumption. Similarly, it processes Excel sheets to execute scripts per user, excluding managerial roles, demonstrating data handling prowess.

These cases underscore repeatability without elaborate guidance. Barry emphasizes commencing with explicit plans, segmenting tasks, and verifying outputs. For Git repositories, Claude clones projects, inspects commit histories, and pinpoints version-specific issues. In Kubernetes contexts, it traverses namespaces, scrutinizes deployments, and peruses pod logs expeditiously.

However, executions demand vigilance. Barry recounts an episode where Claude rebooted a machine prematurely, failing to update boot configurations correctly, underscoring the imperative for output scrutiny. Nonetheless, the tool’s self-correction upon feedback enhances reliability.

Evaluation of Outcomes and Derived Insights

Assessing these applications reveals both efficacies and deficiencies. Successes include adept repository analysis, where Claude discerned alterations across versions, aiding troubleshooting. Its proficiency in interlinking configurations—such as database credentials in application properties—proves invaluable for intricate setups. Moreover, it accelerates tool acquisition, beneficial for client engagements involving novel technologies.

In Kubernetes diagnostics, Claude’s rapid log inspection outpaces manual efforts, facilitating swift resolutions. Log analysis on sanitized files identifies anomalies effectively, while test data generation populates schemas comprehensively. One-off automations address procrastinated tasks, and local container setups streamline development without advanced frameworks.

Conversely, pitfalls abound. Premature completion declarations necessitate clear doneness criteria and measurable objectives. Reading comprehension lapses, as in the misinterpretation of grub update outputs, mimic human errors but require intervention. Context exhaustion precipitates erratic behavior, mandating task fragmentation.

Barry advises defining scopes meticulously, verifying successes, and managing contexts to avert spirals. Despite these, the tool’s utility in DevOps outweighs risks when confined to non-production realms.

Ramifications and Prospective Trajectories

The implications extend to redefining DevOps workflows, where AI aides like Claude diminish manual toil, permitting focus on strategic endeavors. This fosters agility, particularly in compliance and reporting, where generated artifacts ensure regulatory adherence efficiently.

Looking ahead, the convergence of open-source models like Mistral with frontier capabilities portends broader accessibility. Barry speculates that simpler deployments may soon operate on local models, reducing dependency on proprietary services. Tools like Aider, permitting model selection, herald this shift.

In essence, Claude Code’s repurposing exemplifies AI’s potential in operational spheres, promoting efficiency while necessitating prudent governance. As models evolve, their integration into daily practices promises transformative, albeit cautious, advancements in technology management.

Links:

  • Lecture video: https://www.youtube.com/watch?v=nPoC6m3axeU
  • Barry van Someren on LinkedIn: https://www.linkedin.com/in/barryvansomeren
  • Barry van Someren on Twitter/X: https://twitter.com/bvansomeren
  • CoffeeSprout ICT Services website: https://www.coffeesprout.nl/

PostHeaderIcon [VoxxedDaysTicino2026] May the Control Plane Be with You: Kamaji and the Rise of Kubernetes at Scale

Lecturer

Dario Tranchitella serves as the Chief Technology Officer at Clastix, a startup he co-founded in 2020 during the global pandemic. With a background as a site reliability engineer and software developer, Dario specializes in Kubernetes engineering and multi-tenancy solutions. He has extensive experience managing large-scale Kubernetes fleets and contributes to open-source projects, drawing from his prior roles in the tech industry. Relevant links include his LinkedIn profile (https://it.linkedin.com/in/dariotranchitella) and Clastix’s website (https://clastix.io/).

Abstract

This article explores Dario Tranchitella’s insights into scaling Kubernetes through Kamaji, an open-source initiative transforming Kubernetes into a control-plane-as-a-service platform. Originating from real operational challenges, the discussion dissects Kubernetes architecture, the hosted control plane model, community-driven evolution, and adoption by major entities. It analyzes methodologies for multi-tenancy, resource optimization, and resilience, while considering implications for large-scale deployments in cloud-native environments.

Origins and Challenges in Kubernetes Management

Dario’s journey with Kamaji began amid personal and professional turmoil, exemplified by an outage during his father’s wedding that required restoring a Kubernetes cluster. This incident underscored the operational and financial hurdles of scaling Kubernetes beyond a few clusters. As a former site reliability engineer managing a fleet for a U.S. company, Dario encountered the complexities of multi-tenancy, where infrastructure or applications are shared among tenants—be they customers or internal teams—while ensuring fair resource allocation and preventing privilege escalation.

Kubernetes, donated to the Cloud Native Computing Foundation (CNCF), orchestrates containers in a distributed system comprising a control plane and worker nodes. The control plane acts as the “brain,” maintaining application states, while worker nodes provide computational power. Dario likens this to a reconciliation loop: users specify desired states, and Kubernetes aligns current states accordingly, handling tasks like load balancing without manual intervention. It runs ubiquitously—on laptops, clouds, bare metal, or edge devices—abstracting deployment details.

However, scaling introduces bottlenecks. The control plane includes the API server for information handling, the controller manager for reconciliation loops, the scheduler for pod placement to avoid single points of failure, and etcd for state storage using the Raft consensus algorithm. Etcd requires at least three instances for fault tolerance (n/2 + 1), making it resource-intensive and a primary challenge in multi-tenant setups.

In multi-tenancy, Dario emphasizes dividing resources imperatively, akin to apartments in a building: tenants occupy their spaces without infringing on others. Kubernetes excels here, but traditional setups demand separate clusters per tenant to isolate workloads, leading to overhead. Dario’s prior experience revealed inefficiencies, prompting Kamaji’s creation to address these pain points.

The Kamaji Architecture and Hosted Control Plane Model

Kamaji redefines Kubernetes by running control planes as regular pods within a management cluster, adopting a hosted control plane architecture. This separates control planes from worker nodes, allowing a single management cluster to host multiple tenant control planes efficiently. Worker nodes join via the management cluster’s API endpoint, optimizing resources and reducing costs.

Dario contrasts this with traditional setups: instead of dedicating machines per control plane, Kamaji leverages Kubernetes’ scheduling for etcd and other components as pods. This “Kubernetes-in-Kubernetes” approach, inspired by Google’s 2017 Kubernetes Engine, avoids vendor lock-in by supporting tools like kubeadm for certificate management and cluster bootstrapping.

Key innovations include multi-tenant datastores: Kamaji supports etcd, PostgreSQL, or MySQL, allowing collision of databases into single instances for optimization, though Dario advises multiple clusters to minimize blast radius. Scalability tests show a single management cluster handling up to a thousand control planes, but he recommends diversification for resilience.

Methodologically, Kamaji integrates with community projects like Cluster API for node provisioning across providers (Azure, AWS, Google). It avoids reinventing orchestration, focusing solely on control planes while enabling seamless worker node integration. Code samples illustrate simplicity:

apiVersion: kamaji.clastix.io/v1alpha1
kind: TenantControlPlane
metadata:
  name: example
spec:
  kubernetes:
    version: v1.25.0
  dataStore:
    name: default

This YAML defines a tenant control plane, specifying Kubernetes version and datastore, demonstrating declarative management.

Implications include cost savings—reducing dedicated machines—and operational ease, as upgrades affect only the management cluster without tenant disruption.

Community Collaboration and Evolution of Kamaji

Kamaji’s growth stems from open-source collaboration since its 2022 launch at KubeCon Valencia. Dario highlights cross-pollination with organizations like NVIDIA, Rackspace, OVH, Ionos, and the CNCF community. Early adopters provided feedback, debunking scalability myths and proving PostgreSQL viability as an etcd alternative.

Dario’s philosophy: “Do what you love,” drove pursuits like running Kubernetes on PostgreSQL, challenging skeptics. Community tools like Kine (etcd shim) enabled alternative datastores, enhancing flexibility.

Evangelism involved panels at conferences, demystifying hosted control planes alongside Red Hat’s Hypershift and Mirantis’ K0s. Despite similarities, Kamaji’s vanilla Kubernetes focus and multi-datastore support differentiate it.

Code integration with kubeadm ensures portability:

kamaji create --kubeadm-config /path/to/config.yaml

This command bootstraps clusters, allowing imports from existing setups without lock-in.

Consequences: Kamaji fosters a collaborative ecosystem, reducing proprietary dependencies and promoting standards. Adoption by giants validates its scalability, though Dario cautions against over-reliance on single clusters.

Implications for Cloud-Native Scalability and Future Directions

Kamaji addresses Kubernetes’ scaling pains by commoditizing control planes, lowering barriers for multi-tenant platforms. It optimizes resources, crucial in cloud environments where costs accumulate. By hosting control planes as pods, it leverages Kubernetes’ strengths for self-management, a meta-approach enhancing resilience.

Broader implications include democratizing large-scale deployments: smaller teams manage vast fleets without proportional infrastructure. However, Dario stresses evaluating trade-offs—colliding datastores risks contention, necessitating careful architecture.

Future directions involve deeper community integration, potentially expanding to more datastores or advanced scheduling. Kamaji’s open-source ethos ensures evolution through contributions, avoiding silos.

In conclusion, Dario’s work with Kamaji exemplifies pragmatic innovation in cloud-native computing, balancing efficiency, resilience, and community-driven progress.

Links:

PostHeaderIcon [MiamiJUG] Bridging the Gap: A Java Developer’s Guide to the Go Ecosystem

Lecturer

Vladimir Vivien is a veteran software engineer with over 20 years of experience in the technology industry. A specialist in distributed systems and cloud-native architecture, Vladimir spent the first decade of his career as a dedicated Java developer before transitioning to the Go programming language roughly twelve years ago. He is the author of the authoritative text Learning Go Programming and the creator of the LinkedIn Learning course Programming with Go Modules. Vladimir is a passionate advocate for well-architected solutions and currently focuses on building high-performance systems that leverage Go’s unique concurrency primitives.

Abstract

As the backbone of cloud-native infrastructure, the Go programming language (Golang) has become an essential tool for modern software engineering. This article provides a comparative analysis of Go and Java, designed specifically for practitioners familiar with the Java Virtual Machine (JVM) ecosystem. While both languages share a commitment to static typing and garbage collection, they diverge significantly in their approaches to concurrency, deployment, and error handling. By exploring Go’s syntax, its “share by communicating” philosophy via channels, and its deterministic build system, this study highlights how Go simplifies common programming tasks while maintaining the performance required for large-scale systems like Kubernetes and Docker. The analysis concludes by examining Go’s role in the industry and its strategic advantages for distributed architectures.

The Origins and Industry Adoption of Go

Go was developed at Google to solve large-scale software engineering challenges. It was designed not merely as a language, but as a comprehensive suite of tools to address issues like packaging, supply chain security, and build-time performance. Since its public release in 2009, Go has consistently ranked among the most loved languages by developers.

Go’s dominance is particularly evident in the cloud-native and DevOps sectors. Critical infrastructure tools such as Kubernetes, Docker, Terraform, and Prometheus are all written in Go. This is not coincidental; Go’s ability to compile into a single, static binary with fast startup times and low memory overhead makes it ideal for containerized environments. Vladimir notes that while Java offers “Write Once, Run Anywhere” via the JVM, Go provides “Write Once, Compile Anywhere,” targeting specific architectures with a highly optimized toolchain.

Comparative Architecture: Go vs. Java

For the Java developer, Go introduces several paradigm shifts in how code is structured and executed:

Static Typing and Inference

Both languages utilize strict static type systems. However, Go supports implicit typing through the := short variable declaration operator, allowing the compiler to infer the type based on the assigned value. This provides the brevity of a dynamic language while maintaining the safety of static checks at compile time.

Garbage Collection

Go and Java are both garbage-collected. However, whereas Java provides developers with numerous “knobs” and parameters to tune the JVM’s garbage collector, Go takes a minimalist approach. The Go runtime is designed to deliver sub-millisecond GC pauses with almost no manual configuration, relying on compiler optimizations and escape analysis to manage memory efficiently.

Concurrency: Go-routines and Channels

The most significant departure from Java’s threading model is Go’s approach to concurrency. Instead of heavy OS-level threads, Go uses “go-routines”—lightweight threads managed by the Go runtime that cost only a few kilobytes of memory.

Go’s philosophy of concurrency is summarized as: “Do not communicate by sharing memory; instead, share memory by communicating.” This is achieved through Channels, conduits that allow go-routines to pass data safely without the need for traditional locks or race condition worries.

Example of a basic worker pattern in Go:

func worker(id int, jobs <-chan int, results chan<- int) {
    for j := range jobs {
        results <- j * 2
    }
}

func main() {
    jobs := make(chan int, 100)
    results := make(chan int, 100)

    for w := 1; w <= 3; w++ {
        go worker(w, jobs, results) // Launch 3 lightweight go-routines
    }

    for j := 1; j <= 5; j++ {
        jobs <- j
    }
    close(jobs)
    // Results are popped out as they are processed
}

Explicit Error Handling and Resource Management

Unlike Java, which relies on a hierarchy of Exceptions that bubble up the call stack, Go requires explicit error handling. Functions in Go can return multiple values, and by convention, the last value is often an error type.

Vladimir explains that this “check everything” approach prevents silent failures and forces developers to consider failure states as part of the primary logic flow. Additionally, Go replaces Java’s try-with-resources or finally blocks with the defer keyword, which schedules a function call (like closing a file or network connection) to run immediately before the surrounding function returns.

Conclusion: Where Go Shines

Go’s design choices prioritize simplicity, readability, and performance. It excels in building CLI tools, distributed systems, and high-performance APIs capable of handling thousands of concurrent connections out of the box. For the Java developer, Go offers a streamlined alternative that reduces the complexity of modern cloud-native development without sacrificing the robustness required for enterprise-scale engineering.

Links:

PostHeaderIcon [DevoxxGR2025] Optimized Kubernetes Scaling with Karpenter

Alex König, an AWS expert, delivered a 39-minute talk at Devoxx Greece 2025, exploring how Karpenter enhances Kubernetes cluster autoscaling for speed, cost-efficiency, and availability.

Karpenter’s Dynamic Autoscaling

König introduced Karpenter as an open-source, Kubernetes-native autoscaling solution, contrasting it with the traditional Cluster Autoscaler. Unlike the latter, which relies on uniform node groups (e.g., nodes with four CPUs and 16GB RAM), Karpenter uses the EC2 Fleet API to dynamically provision nodes tailored to workload needs. For instance, if a pod requires one CPU, Karpenter allocates a node with minimal excess capacity, avoiding resource waste. This right-sizing, combined with groupless scaling, enables faster and more cost-effective scaling, especially in dynamic environments.

Ensuring Availability with Constraints

König addressed availability challenges reported by users, emphasizing Kubernetes-native scheduling constraints to mitigate disruptions. Topology spread constraints distribute pods across availability zones, reducing the risk of downtime if a node fails. Pod disruption budgets, affinity/anti-affinity rules, and priority classes further ensure critical workloads are scheduled appropriately. For stateful workloads using EBS, König recommended setting the volume binding mode to “wait for first consumer” to avoid pod-volume mismatches across zones, preventing crashes and ensuring reliability.

Integrating with KEDA for Application Scaling

For advanced scaling, König highlighted combining Karpenter with KEDA for event-driven, application-specific scaling. KEDA scales pods based on metrics like Kafka topic sizes or SQS queues, beyond CPU/memory. Karpenter then provisions nodes for pending pods, enabling seamless scaling for workloads like flash sales. König outlined a four-step migration from Cluster Autoscaler to Karpenter, emphasizing its simplicity and open-source documentation.

Links

PostHeaderIcon [GopherConUK2025] CPU Quota Semantics and Runtime Scheduler Behavior in Containerized Environments

Lecturer

Bill Kennedy is a software engineer, technical trainer, and Managing Partner at Ardan Labs. He has authored multiple technical books on Go programming and serves as a core organizer for developer communities worldwide. His professional work focuses on high-performance backend development, system design, and training software engineering teams on runtime internals and concurrent programming semantics.

Abstract

Deploying managed language runtimes into containerized orchestration frameworks requires a comprehensive understanding of how compute limits interact with application-level scheduling primitives. This article examines the behavior of the Go runtime scheduler when executed under Kubernetes CPU limits. By analyzing thread management, operating system context switching, and the mechanics of Completely Fair Scheduler (CFS) quota enforcement, this study highlights performance degradation scenarios caused by misalignment between thread allocation and container CPU constraints. Furthermore, empirically derived benchmarking demonstrates how adjusting runtime concurrency configurations mitigates kernel-level throttling and improves request throughput in CPU-bound and IO-bound application workloads.

Micro-Architecture, Concurrency, and Context Switching Mechanics

Modern multi-core processors execute operations via clock cycles, where instruction execution frequency depends on pipelined hardware architectures. On a standard processor core running at a 3 GHz clock rate, a single nanosecond corresponds to three clock cycles. Leveraging superscalar execution pipelines, modern hardware can process up to four instructions per clock cycle on average, yielding approximately twelve instructions per nanosecond. Consequently, operational latencies—whether originating from memory access, network round trips, or kernel thread context switches—directly translate into unexecuted instruction cycles.

Operational Event Approximate Duration Lost Instruction Opportunities
OS Thread Context Switch 1,000 ns (1 µs) ~12,000 instructions
Datacenter Network Round Trip 500,000 ns (0.5 ms) ~6,000,000 instructions
Go Routine Context Switch 200 ns ~2,400 instructions

In system software, workloads are categorized as either CPU-bound or IO-bound. CPU-bound tasks execute uninterrupted mathematical or logical operations, utilizing their full operating system time slice. Under CPU-bound conditions, context switches incur overhead that degrades throughput unless application thread counts strictly match available physical cores. Conversely, IO-bound workloads frequently transition threads into blocked or waiting states due to asynchronous network calls or file interactions.

The Go runtime abstracts operating system (OS) threads through an M:N scheduler, mapping M goroutines (application-level lightweight threads) onto N OS threads managed across logical processors known as P structures. Physical CPU cores are abstracted into these P units, which hold local run queues for goroutines. The Go scheduler operates as a work-stealing system: idle P structures steal runnable goroutines from other local queues or a global run queue.

+-----------------------------------------------+
|                 OS Kernel                     |
|  [Core 0]    [Core 1]    [Core 2]    [Core 3] |
+-----------------------------------------------+
       ^          ^           ^           ^
       |          |           |           |
    [  M  ]    [  M  ]     [  M  ]     [  M  ]
       |          |           |           |
    [  P  ]    [  P  ]     [  P  ]     [  P  ]
    /     \      ...         ...         ...
 [ G ]   [ G ]

To maximize thread utilization, asynchronous system calls (such as network operations) are handled via a dedicated network poller thread. When a goroutine initiates a network read, the runtime detaches the goroutine from its current M thread and registers it with the network poller. This frees the underlying M thread to immediately execute other goroutines assigned to that logical P processor. Synchronous operations, such as blocking file system IO, force the runtime to decouple the blocking M thread from its assigned P structure and allocate or unpark a separate OS thread to keep the P processor active.

Through this abstraction, the Go runtime transforms application-level IO-bound tasks into CPU-bound operational streams from the operating system’s perspective. The OS kernel observes saturated worker threads (M), allowing them to consume allocated time slices efficiently without premature thread parking.

Kubernetes Completely Fair Scheduler (CFS) Quota Semantics

Kubernetes enforces compute resource limits using Linux control groups (cgroups) via the Completely Fair Scheduler (CFS) quota system. A CPU resource limit specified in millicores (such as 250m) translates into a time-based allocation per enforcement period. By default, the Linux kernel CFS operates on a 100-millisecond period.

Allocated Time Formula:
Allocated Time = CFS Period * (Millicores / 1000)

For an allocation of 250m across a 100 ms period, the container receives exactly 25 ms of cumulative execution time:

Allocated Time = 100 ms * (250 / 1000) = 25 ms

Crucially, the Linux CFS tracks CPU quota consumption cumulatively across all running OS threads within the container’s thread group. If an application spawns 16 threads that execute concurrently on a multi-core host system, each thread consumes physical core time simultaneously.

Quota Exhaustion Rate:
Quota Exhaustion Rate = Number of Threads * Elapsed Time

With 16 active OS threads, a 25 ms CPU quota is depleted in less than 2 ms of real time:

Exhaustion Time = 25 ms / 16 = 1.5625 ms

Once the total execution time across all threads reaches the 25 ms ceiling, the kernel CFS throttles the entire container. The container processes remain paused until the 100 ms cycle resets, resulting in severe latency spikes and degraded service throughput.

Architectural Misalignment: Go Max Procs in Container Runtimes

By default, the Go runtime initializes the number of logical processors (P) via the GOMAXPROCS variable based on system calls that query host core availability. In standard Kubernetes pod deployments without explicit runtime tuning, the runtime inspects the host node rather than container cgroup boundaries.

If a pod configured with a 250m limit is scheduled on a 16-core physical node, GOMAXPROCS defaults to 16. The runtime creates 16 logical P processors and corresponding OS threads (M).

Container Configuration: CPU Limit = 250m (25ms per 100ms cycle)
Host Infrastructure: 16 Physical Cores
Default Go Runtime Behavior: GOMAXPROCS = 16

+-------------------------------------------------------+
| 16 OS Threads (M) Executing Simultaneously           |
| [M1] [M2] [M3] [M4] [M5] [M6] ... [M16]               |
+-------------------------------------------------------+
                           |
                           v
    Consumes 25ms Quota in ~1.56ms of Real Time
                           |
                           v
+-------------------------------------------------------+
| Kernel CFS Throttles Container for Remaining ~98.4ms  |
+-------------------------------------------------------+

When incoming requests hit the container, all 16 worker threads wake up to process goroutines. The cumulative CPU time consumed by these 16 concurrent threads exhausts the 25 ms cgroup quota almost instantly. The application spends the vast majority of every 100 ms enforcement window in a kernel-throttled state.

To resolve this misalignment, the application runtime must match its logical thread capacity to its cgroup quota boundaries. Setting GOMAXPROCS=1 forces the Go scheduler to utilize a single logical P processor and one primary operating thread, executing sequential instructions over the full 25 ms window without premature multi-threaded quota depletion.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: sales-service
spec:
  template:
    spec:
      containers:
      - name: service
        image: sales-service:1.0
        env:
        - name: GOMAXPROCS
          valueFrom:
            resourceFieldRef:
              resource: limits.cpu
        resources:
          limits:
            cpu: "250m"

In Go deployments, setting GOMAXPROCS via container environment variables applies a mathematical ceiling function to convert fractional core limits into discrete thread bounds.

Experimental Evaluation and System Optimization

Empirical performance tests were conducted on a Kubernetes cluster managed via kind hosted on an 16-core machine. The microservice stack comprised an HTTP API service backed by a PostgreSQL database. Load testing was executed using automated benchmark tools transmitting synthetic HTTP workloads.

Test Configuration A: Default Core Allocation

  • CPU Limit: 250m (25 ms per 100 ms)
  • Host Cores Detected: 16
  • GOMAXPROCS: 16 (Default)

Test Configuration B: Matched Runtime Constraints

  • CPU Limit: 250m (25 ms per 100 ms)
  • Host Cores Detected: 16
  • GOMAXPROCS: 1 (Explicitly configured)

Measured Experimental Results

Metric Config A (GOMAXPROCS=16) Config B (GOMAXPROCS=1) Performance Impact
Throughput ~126 req/sec ~2,746 req/sec ~21.7x Increase
P99 Latency ~200 ms ~3.6 ms ~98.2% Reduction

Constraining the runtime thread count to align with container limits produced a 21-fold throughput increase while eliminating excessive tail latency caused by kernel CFS throttling.

Cascading Latency and Upstream Service Dependencies

In distributed microservice topologies, runtime throttling can cascade across service boundaries. During secondary experimentation, an authentication service dependency (auth-service) was assigned a restricted CPU limit (100m).

Even when the primary edge service (sales-service) was provisioned with unrestricted CPU allocations, overall request throughput dropped to baseline throttled levels. Blocking latencies introduced by the throttled upstream dependency bottlenecked the unconstrained downstream service. Diagnosing performance anomalies requires evaluating total system execution graphs rather than isolating individual application metrics.

Links: