Recent Posts
Archives

Posts Tagged ‘CloudNative’

PostHeaderIcon [VoxxedDaysBucharest2026] Optimizing LLM Inference on Kubernetes: Abdel Sghiouar on Practical Techniques for the Rest of Us

Lecturer

Abdel Sghiouar is a Developer Advocate at Google Cloud with deep expertise in cloud-native technologies, Kubernetes orchestration, and AI/ML workload optimization. Drawing from a robust background in infrastructure engineering and open source contributions, Abdel helps organizations design, deploy, and tune complex AI applications for production environments across diverse infrastructures.

Abstract

While major cloud providers and hyperscalers leverage virtually unlimited computational resources, the majority of organizations face significant constraints when operationalizing Large Language Models. Abdel Sghiouar presents a comprehensive set of practical strategies for optimizing LLM inference workloads on Kubernetes. The session systematically addresses container and model optimization techniques, accelerator management, data persistence and storage considerations, networking and intelligent load balancing, and advanced observability practices. Emphasis is placed on open-source tools and architectural patterns that deliver meaningful cost-performance improvements adaptable to on-premises, hybrid, and public cloud deployments.

Understanding LLM Inference Characteristics and Challenges

Large Language Models continue their rapid evolution in both scale and sophistication. Architectural innovations such as mixture-of-experts (MoE) enable dynamic activation of specialized sub-networks, while multi-modal capabilities process diverse inputs including text, images, audio, and video. Expanded context windows support richer interactions but demand substantial memory resources.

Inference execution comprises two primary phases with contrasting characteristics: the prefill stage (encoding input tokens, predominantly compute-bound) and the decode stage (token generation, typically memory-bound). KV (key-value) caching optimizes conversational flows by preserving intermediate states, avoiding redundant prefill computations for subsequent messages.

Deployment topologies vary considerably. Single-host single-accelerator setups predominate for local development and experimentation (e.g., using Ollama). Single-host multi-accelerator configurations require model sharding across GPUs within one machine. Multi-host distributed deployments introduce complex requirements for high-bandwidth, low-latency interconnects to maintain coherent context across nodes. Each topology presents distinct challenges regarding scalability, fault tolerance, and operational complexity.

Container, Model, and Storage Optimizations

Inference serving runtimes and model artifacts generate exceptionally large container images, frequently exceeding several gigabytes prior to incorporating weights. Conventional optimization strategies like multi-stage builds or native compilation (e.g., GraalVM) prove inadequate for these workloads.

Distributed caching solutions such as Spiegel provide cluster-wide image and model artifact caching, substantially reducing repeated pulls from external registries. Kubernetes-native features enabling containers as volumes allow separate packaging of models, which can then be mounted efficiently onto serving runtimes. When combined with caching layers, these approaches dramatically accelerate cold starts.

Quantization techniques offer another lever, reducing numerical precision (e.g., FP16 to INT8 or lower) to decrease memory footprints while preserving sufficient accuracy for many applications. Careful selection of quantization levels based on task sensitivity balances performance and quality.

Accelerator Management and Dynamic Resource Allocation

Kubernetes has supported GPU scheduling through device plugins for several years. However, static device configurations struggle with real-world constraints including accelerator scarcity and heterogeneous hardware fleets.

Dynamic Resource Allocation, matured in recent Kubernetes versions, introduces flexible resource claiming based on abstract characteristics rather than rigid device specifications (e.g., requesting “NVIDIA GPU with minimum 30GB memory and specific core count”). This enables more efficient scheduling across mixed clusters and better utilization rates.

Integration with cluster autoscalers allows on-demand provisioning, addressing both availability gaps and cost optimization by scaling resources precisely to workload demands. Platform operators describe device inventories; application teams specify requirements, with the scheduler performing intelligent matching.

Networking, Load Balancing, and Observability Considerations

LLM traffic profiles differ markedly from conventional web workloads. Requests exhibit high variability in size and computational intensity (simple text queries versus multi-modal inputs), while responses frequently involve streaming token generation. Standard round-robin load balancing produces inefficient distributions, with certain backends becoming overloaded while others remain underutilized.

The Kubernetes Gateway API, augmented with custom endpoint selection logic, supports sophisticated routing decisions based on request attributes extracted from bodies (model identifier, input modality, streaming requirements) combined with real-time backend telemetry. This facilitates intelligent traffic steering, prioritization of business-critical workloads, and maintenance of sticky sessions necessary for coherent streaming interactions.

Comprehensive observability must encompass prefill and decode phase latencies, KV cache hit rates, token generation throughput, GPU utilization, and end-to-end request metrics. Integration with Prometheus, Grafana, and specialized LLM monitoring solutions provides actionable insights for capacity planning and bottleneck identification.

Practical Patterns and the LLM-D Project

The LLM-D initiative, hosted under the Linux Foundation with contributions from Google, IBM, NVIDIA, and additional partners, aggregates architectural patterns, performance benchmarks, and reference implementations for production-grade inference. Key elements include optimized prefill/decode separation, advanced routing logic often leveraging engines like vLLM, and comprehensive guidance for multi-node deployments.

A holistic, layered optimization strategy proves most effective: infrastructure-level improvements (caching, persistent volumes), platform capabilities (dynamic scheduling, intelligent networking), and application-level choices (model quantization, serving engine selection). Organizations without hyperscale resources can still achieve competitive efficiency and scalability through disciplined application of these patterns.

Links:

PostHeaderIcon [DevoxxGR2026] GenAI on Kubernetes: Training, Inference, and Serving in Production Environments

Lecturer
Alessandro Vozza is a seasoned cloud-native advocate and technologist with deep expertise in Kubernetes and AI/ML operations. He contributes actively to open-source communities and focuses on practical, scalable deployments of generative AI workloads. As a speaker and practitioner, Alessandro emphasizes operational excellence, resource efficiency, and the integration of modern AI tools within established cloud-native platforms.

Abstract
In this hands-on tutorial at Devoxx Greece 2026, Alessandro Vozza guides developers through the complete lifecycle of running generative AI workloads on Kubernetes. From distributed training jobs with GPU scheduling to optimized inference and scalable model serving, the session demonstrates how to leverage operators, autoscaling, vector stores, and frameworks like KServe, Ray, vLLM, and Kubeflow. Attendees gain actionable insights into designing efficient GPU clusters, fine-tuning models securely, and deploying production-grade architectures that integrate seamlessly with existing Kubernetes expertise.

The Convergence of Kubernetes and Generative AI

Kubernetes has evolved into the de facto platform for orchestrating complex, resource-intensive workloads, including those powered by generative AI. Vozza begins by contextualizing the challenges: training large models demands massive parallel computation across GPUs, inference requires low-latency serving under variable traffic, and the entire pipeline must remain observable, secure, and cost-effective. Traditional approaches struggle with these demands, but Kubernetes patterns—scheduling, autoscaling, and declarative resource management—provide a robust foundation.

The session highlights how the community has responded with specialized tools. Projects like Kubeflow address the full ML lifecycle, while KServe and vLLM focus on high-performance inference. These build upon core Kubernetes capabilities, allowing teams to treat AI workloads with the same rigor applied to microservices.

Distributed Training and GPU Orchestration

Training generative models is computationally intensive and benefits enormously from Kubernetes’ scheduling strengths. Vozza demonstrates launching distributed training jobs, emphasizing GPU-aware scheduling through device plugins and resource requests. Nodes are labeled with GPU capacity, enabling the scheduler to place pods on suitable hardware.

The tutorial covers hyperparameter tuning with tools like Katib, which automates experimentation across multiple configurations. Fine-tuning involves augmenting base models with domain-specific data, a process that Kubernetes orchestrates reliably through persistent volumes and checkpointing. Attendees learn to monitor training progress using built-in observability and handle failures gracefully with retries and job controllers.

Resource efficiency emerges as a key theme. Techniques such as multi-instance GPU (MIG) partitioning allow a single physical GPU to support multiple smaller workloads, maximizing utilization without over-provisioning expensive hardware.

Inference Serving and Model Deployment

Once trained, models must be served efficiently. Vozza walks through deploying inference endpoints with KServe, which abstracts the complexities of scaling and routing. vLLM serves as the high-throughput inference engine, leveraging continuous batching and paged attention for superior performance.

The architecture supports multi-model serving, where a single deployment handles various models based on request characteristics. Gateway API extensions make the ingress layer LLM-aware, enabling intelligent routing based on factors like key-value cache state or model specialization. This ensures optimal resource allocation and minimal latency.

Autoscaling plays a critical role. Horizontal Pod Autoscaler (HPA) combined with KEDA reacts to custom metrics such as queue depth or tokens processed per second, dynamically adjusting replicas to match demand while controlling costs.

Operational Considerations and Best Practices

Production readiness demands comprehensive observability. Vozza integrates Prometheus exporters and logging to track token throughput, latency, and GPU utilization. Security best practices include least-privilege access for model endpoints and encrypted communication.

The tutorial addresses common pitfalls: managing model registries for versioning, handling cold starts through caching, and ensuring reproducibility across environments. By treating models as first-class Kubernetes citizens, teams achieve consistent deployments from development to production.

Practical Roadmap and Future Directions

Participants receive a working reference setup they can adapt immediately. Vozza encourages starting small—perhaps with a single-model inference service—before scaling to distributed training and multi-model architectures. The session reinforces that Kubernetes knowledge directly transfers to AI operations, lowering the barrier for traditional platform teams.

Looking ahead, evolving features like dynamic resource allocation and improved GPU topology awareness will further streamline GenAI workloads. The message is clear: Kubernetes is not merely compatible with generative AI; it is becoming the preferred operational layer for the entire lifecycle.

Links:

PostHeaderIcon [VoxxedDaysTicino2026] May the Control Plane Be with You: Kamaji and the Rise of Kubernetes at Scale

Lecturer

Dario Tranchitella serves as the Chief Technology Officer at Clastix, a startup he co-founded in 2020 during the global pandemic. With a background as a site reliability engineer and software developer, Dario specializes in Kubernetes engineering and multi-tenancy solutions. He has extensive experience managing large-scale Kubernetes fleets and contributes to open-source projects, drawing from his prior roles in the tech industry. Relevant links include his LinkedIn profile (https://it.linkedin.com/in/dariotranchitella) and Clastix’s website (https://clastix.io/).

Abstract

This article explores Dario Tranchitella’s insights into scaling Kubernetes through Kamaji, an open-source initiative transforming Kubernetes into a control-plane-as-a-service platform. Originating from real operational challenges, the discussion dissects Kubernetes architecture, the hosted control plane model, community-driven evolution, and adoption by major entities. It analyzes methodologies for multi-tenancy, resource optimization, and resilience, while considering implications for large-scale deployments in cloud-native environments.

Origins and Challenges in Kubernetes Management

Dario’s journey with Kamaji began amid personal and professional turmoil, exemplified by an outage during his father’s wedding that required restoring a Kubernetes cluster. This incident underscored the operational and financial hurdles of scaling Kubernetes beyond a few clusters. As a former site reliability engineer managing a fleet for a U.S. company, Dario encountered the complexities of multi-tenancy, where infrastructure or applications are shared among tenants—be they customers or internal teams—while ensuring fair resource allocation and preventing privilege escalation.

Kubernetes, donated to the Cloud Native Computing Foundation (CNCF), orchestrates containers in a distributed system comprising a control plane and worker nodes. The control plane acts as the “brain,” maintaining application states, while worker nodes provide computational power. Dario likens this to a reconciliation loop: users specify desired states, and Kubernetes aligns current states accordingly, handling tasks like load balancing without manual intervention. It runs ubiquitously—on laptops, clouds, bare metal, or edge devices—abstracting deployment details.

However, scaling introduces bottlenecks. The control plane includes the API server for information handling, the controller manager for reconciliation loops, the scheduler for pod placement to avoid single points of failure, and etcd for state storage using the Raft consensus algorithm. Etcd requires at least three instances for fault tolerance (n/2 + 1), making it resource-intensive and a primary challenge in multi-tenant setups.

In multi-tenancy, Dario emphasizes dividing resources imperatively, akin to apartments in a building: tenants occupy their spaces without infringing on others. Kubernetes excels here, but traditional setups demand separate clusters per tenant to isolate workloads, leading to overhead. Dario’s prior experience revealed inefficiencies, prompting Kamaji’s creation to address these pain points.

The Kamaji Architecture and Hosted Control Plane Model

Kamaji redefines Kubernetes by running control planes as regular pods within a management cluster, adopting a hosted control plane architecture. This separates control planes from worker nodes, allowing a single management cluster to host multiple tenant control planes efficiently. Worker nodes join via the management cluster’s API endpoint, optimizing resources and reducing costs.

Dario contrasts this with traditional setups: instead of dedicating machines per control plane, Kamaji leverages Kubernetes’ scheduling for etcd and other components as pods. This “Kubernetes-in-Kubernetes” approach, inspired by Google’s 2017 Kubernetes Engine, avoids vendor lock-in by supporting tools like kubeadm for certificate management and cluster bootstrapping.

Key innovations include multi-tenant datastores: Kamaji supports etcd, PostgreSQL, or MySQL, allowing collision of databases into single instances for optimization, though Dario advises multiple clusters to minimize blast radius. Scalability tests show a single management cluster handling up to a thousand control planes, but he recommends diversification for resilience.

Methodologically, Kamaji integrates with community projects like Cluster API for node provisioning across providers (Azure, AWS, Google). It avoids reinventing orchestration, focusing solely on control planes while enabling seamless worker node integration. Code samples illustrate simplicity:

apiVersion: kamaji.clastix.io/v1alpha1
kind: TenantControlPlane
metadata:
  name: example
spec:
  kubernetes:
    version: v1.25.0
  dataStore:
    name: default

This YAML defines a tenant control plane, specifying Kubernetes version and datastore, demonstrating declarative management.

Implications include cost savings—reducing dedicated machines—and operational ease, as upgrades affect only the management cluster without tenant disruption.

Community Collaboration and Evolution of Kamaji

Kamaji’s growth stems from open-source collaboration since its 2022 launch at KubeCon Valencia. Dario highlights cross-pollination with organizations like NVIDIA, Rackspace, OVH, Ionos, and the CNCF community. Early adopters provided feedback, debunking scalability myths and proving PostgreSQL viability as an etcd alternative.

Dario’s philosophy: “Do what you love,” drove pursuits like running Kubernetes on PostgreSQL, challenging skeptics. Community tools like Kine (etcd shim) enabled alternative datastores, enhancing flexibility.

Evangelism involved panels at conferences, demystifying hosted control planes alongside Red Hat’s Hypershift and Mirantis’ K0s. Despite similarities, Kamaji’s vanilla Kubernetes focus and multi-datastore support differentiate it.

Code integration with kubeadm ensures portability:

kamaji create --kubeadm-config /path/to/config.yaml

This command bootstraps clusters, allowing imports from existing setups without lock-in.

Consequences: Kamaji fosters a collaborative ecosystem, reducing proprietary dependencies and promoting standards. Adoption by giants validates its scalability, though Dario cautions against over-reliance on single clusters.

Implications for Cloud-Native Scalability and Future Directions

Kamaji addresses Kubernetes’ scaling pains by commoditizing control planes, lowering barriers for multi-tenant platforms. It optimizes resources, crucial in cloud environments where costs accumulate. By hosting control planes as pods, it leverages Kubernetes’ strengths for self-management, a meta-approach enhancing resilience.

Broader implications include democratizing large-scale deployments: smaller teams manage vast fleets without proportional infrastructure. However, Dario stresses evaluating trade-offs—colliding datastores risks contention, necessitating careful architecture.

Future directions involve deeper community integration, potentially expanding to more datastores or advanced scheduling. Kamaji’s open-source ethos ensures evolution through contributions, avoiding silos.

In conclusion, Dario’s work with Kamaji exemplifies pragmatic innovation in cloud-native computing, balancing efficiency, resilience, and community-driven progress.

Links:

PostHeaderIcon [MiamiJUG] Bridging the Gap: A Java Developer’s Guide to the Go Ecosystem

Lecturer

Vladimir Vivien is a veteran software engineer with over 20 years of experience in the technology industry. A specialist in distributed systems and cloud-native architecture, Vladimir spent the first decade of his career as a dedicated Java developer before transitioning to the Go programming language roughly twelve years ago. He is the author of the authoritative text Learning Go Programming and the creator of the LinkedIn Learning course Programming with Go Modules. Vladimir is a passionate advocate for well-architected solutions and currently focuses on building high-performance systems that leverage Go’s unique concurrency primitives.

Abstract

As the backbone of cloud-native infrastructure, the Go programming language (Golang) has become an essential tool for modern software engineering. This article provides a comparative analysis of Go and Java, designed specifically for practitioners familiar with the Java Virtual Machine (JVM) ecosystem. While both languages share a commitment to static typing and garbage collection, they diverge significantly in their approaches to concurrency, deployment, and error handling. By exploring Go’s syntax, its “share by communicating” philosophy via channels, and its deterministic build system, this study highlights how Go simplifies common programming tasks while maintaining the performance required for large-scale systems like Kubernetes and Docker. The analysis concludes by examining Go’s role in the industry and its strategic advantages for distributed architectures.

The Origins and Industry Adoption of Go

Go was developed at Google to solve large-scale software engineering challenges. It was designed not merely as a language, but as a comprehensive suite of tools to address issues like packaging, supply chain security, and build-time performance. Since its public release in 2009, Go has consistently ranked among the most loved languages by developers.

Go’s dominance is particularly evident in the cloud-native and DevOps sectors. Critical infrastructure tools such as Kubernetes, Docker, Terraform, and Prometheus are all written in Go. This is not coincidental; Go’s ability to compile into a single, static binary with fast startup times and low memory overhead makes it ideal for containerized environments. Vladimir notes that while Java offers “Write Once, Run Anywhere” via the JVM, Go provides “Write Once, Compile Anywhere,” targeting specific architectures with a highly optimized toolchain.

Comparative Architecture: Go vs. Java

For the Java developer, Go introduces several paradigm shifts in how code is structured and executed:

Static Typing and Inference

Both languages utilize strict static type systems. However, Go supports implicit typing through the := short variable declaration operator, allowing the compiler to infer the type based on the assigned value. This provides the brevity of a dynamic language while maintaining the safety of static checks at compile time.

Garbage Collection

Go and Java are both garbage-collected. However, whereas Java provides developers with numerous “knobs” and parameters to tune the JVM’s garbage collector, Go takes a minimalist approach. The Go runtime is designed to deliver sub-millisecond GC pauses with almost no manual configuration, relying on compiler optimizations and escape analysis to manage memory efficiently.

Concurrency: Go-routines and Channels

The most significant departure from Java’s threading model is Go’s approach to concurrency. Instead of heavy OS-level threads, Go uses “go-routines”—lightweight threads managed by the Go runtime that cost only a few kilobytes of memory.

Go’s philosophy of concurrency is summarized as: “Do not communicate by sharing memory; instead, share memory by communicating.” This is achieved through Channels, conduits that allow go-routines to pass data safely without the need for traditional locks or race condition worries.

Example of a basic worker pattern in Go:

func worker(id int, jobs <-chan int, results chan<- int) {
    for j := range jobs {
        results <- j * 2
    }
}

func main() {
    jobs := make(chan int, 100)
    results := make(chan int, 100)

    for w := 1; w <= 3; w++ {
        go worker(w, jobs, results) // Launch 3 lightweight go-routines
    }

    for j := 1; j <= 5; j++ {
        jobs <- j
    }
    close(jobs)
    // Results are popped out as they are processed
}

Explicit Error Handling and Resource Management

Unlike Java, which relies on a hierarchy of Exceptions that bubble up the call stack, Go requires explicit error handling. Functions in Go can return multiple values, and by convention, the last value is often an error type.

Vladimir explains that this “check everything” approach prevents silent failures and forces developers to consider failure states as part of the primary logic flow. Additionally, Go replaces Java’s try-with-resources or finally blocks with the defer keyword, which schedules a function call (like closing a file or network connection) to run immediately before the surrounding function returns.

Conclusion: Where Go Shines

Go’s design choices prioritize simplicity, readability, and performance. It excels in building CLI tools, distributed systems, and high-performance APIs capable of handling thousands of concurrent connections out of the box. For the Java developer, Go offers a streamlined alternative that reduces the complexity of modern cloud-native development without sacrificing the robustness required for enterprise-scale engineering.

Links:

PostHeaderIcon [GopherConUK2025] How Just Eat Uses Tooling to Deploy Go Micro-services in Minutes

Lecturer

Ainsley Clark is a Senior Software Engineer at Just Eat, working within the Jet Connect team. He has been with the organisation for just over two years and maintains a strong professional focus on Go. His work centres on the design and operation of the internal microservice development toolkit that underpins large-scale order and menu processing across hundreds of partner integrations.

Abstract

This article describes the evolution and capabilities of GoKit, the internal microservice development toolkit developed by Just Eat’s Jet Connect team. Confronted with the maintenance burden of hundreds of independently forked services, the team replaced a template-based approach with a centralised code-generation and infrastructure-as-code system. The resulting tool scaffolds services, generates event consumers and producers, provisions cloud resources, produces continuous-integration workflows and emits operational metrics with minimal engineer effort. A concrete pizza-order example illustrates how a fully instrumented, event-driven service can be created and extended in minutes rather than days, allowing engineers to concentrate on business logic rather than infrastructure boilerplate.

From Monolith and Templates to a Centralised Toolkit

Jet Connect began life as Flight, an integration platform founded in 2013. Its purpose is to unify order and menu processing across restaurants, groceries and electronics partners so that a single point-of-sale interaction replaces multiple device-specific payloads. Today the platform serves approximately 731 000 partners across seventeen countries and processes hundreds of millions of orders annually. The Jet Connect engineering group itself comprises roughly fifty people and owns more than one hundred Go microservices.

The journey to that estate began with a PHP monolith that became increasingly difficult to change. Integrations were subsequently extracted into TypeScript services that spoke gRPC to the monolith. A GitHub template repository accelerated the creation of new integrations: an engineer could fork the template and obtain environment files, Helm charts and build workflows in minutes. The approach delivered speed in the early days yet exacted a heavy maintenance cost. A bug fix or standardisation change had to be applied manually to every fork. Inconsistent folder structures and coding patterns made context-switching expensive. End-to-end testing was absent, reducing confidence in releases.

The decision was therefore taken to move the entire integration layer to Go and to replace the template model with a purpose-built toolkit. The requirements were exacting: generate or update a repository in seconds; treat OpenAPI documentation as a first-class artefact; allow an engineer to subscribe to events by writing only a handler; hide infrastructure details such as databases, event buses and object storage; auto-generate continuous-integration and continuous-delivery pipelines; guarantee safe production deployment on merge to main; support capability testing across the full service graph; and provide a consistent local development experience with minimal setup. The resulting system is known internally as GoKit.

Scaffolding, Event Handling and Infrastructure as Code

GoKit meets those requirements through a combination of code generation and a declarative service description. The command gokit new <service-name> presents a short interactive questionnaire: does the service consume events, produce events, expose an HTTP server, require a database or object store? On the basis of the answers it scaffolds a consistent folder structure whose most important directories are cmd/app (containing a generated main) and internal (the sole location for custom business logic). A file named service.json acts as the single source of truth for infrastructure; it is consumed by Terraform templates that provision the necessary cloud resources.

A typical event handler receives an HTTP client, an event-bus producer and the incoming event (typed via reflection performed by GoKit). After performing domain work—calling a partner API, writing analytics, storing a payload—the handler emits a successor event. Registration consists of a single method call that associates the handler with a concrete event type. A subsequent gokit update rewrites service.json, updates generated code and refreshes documentation. The resulting README lists every consumed and produced topic, giving any engineer an immediate, accurate overview of the service’s responsibilities without requiring manual documentation effort.

Because the same service.json also drives resource provisioning, adding a DynamoDB table or an S3 bucket is a matter of inserting a few lines of JSON and re-running the update command. Generated interfaces for read/write operations and object-store upload/download encourage dependency injection and make unit testing straightforward. Should the underlying technology later change (for example from DynamoDB to PostgreSQL), only the implementation behind the interface needs to be altered; every consuming service continues to compile and run unchanged. The same mechanism supports vertical and horizontal scaling declarations—minimum replica counts, CPU and memory class—again expressed as simple keys in service.json.

Continuous Integration, Capability Testing and Observability

All continuous-integration workflows are themselves generated by GoKit. A change to a workflow template is made once in the central repository; every service inherits the update the next time gokit update is run. Capability tests spin up the entire relevant service graph under Docker, inject a correlation identifier into every request and event, and assert final outcomes. The approach scales to capability chains that involve ten or twenty intermediate services and provides high confidence before production deployment. Drift detection prevents accidental divergence: if an engineer edits a generated file without running the update command, continuous integration fails the pull request. Version pinning of GoKit itself ensures that services cannot merge while they lag behind a required toolkit version, producing a natural, incremental migration path across the estate.

Operational metrics appear automatically. Grafana dashboards display event rates, HTTP status codes, latency distributions, DynamoDB read/write performance and S3 operation counts without any additional instrumentation code. The same consistency that simplifies development also simplifies observation. When a product request arrives for analytics storage or an order-ready notification endpoint, the engineer adds a handful of lines to service.json or an OpenAPI specification, runs the update command, implements a short handler, and obtains a fully instrumented, tested and documented service ready for review.

Outcomes and Lessons

In the six years of its existence GoKit has allowed the Jet Connect team to update the entire service estate with a single pull request, to maintain uniform folder structures that eliminate costly context switches, and to keep documentation accurate without manual effort. Engineers spend their time on business problems—partner-specific payloads, analytics requirements, notification flows—rather than on Terraform, Helm or continuous-integration boilerplate. The toolkit has been instrumental in scaling the platform to hundreds of millions of orders while keeping the cognitive load on individual developers manageable.

The most transferable lessons are structural rather than technological. A single source of truth for service shape, aggressive code generation of repetitive artefacts, enforcement of consistency through continuous integration, and the deliberate abstraction of infrastructure behind stable interfaces together convert the chronic maintenance burden of a large microservice estate into a manageable, largely automated background process. The result is that complex event-driven workflows can be designed, implemented, tested and deployed in minutes rather than days, freeing engineering capacity for the differentiated work that actually delivers value to partners and customers.
The pizza-service walkthrough, though simplified for presentation, captures the essential rhythm of day-to-day work. An engineer begins with a short questionnaire, receives a fully structured repository, writes a handful of domain-specific lines, runs an update command, and obtains a service that already possesses continuous-integration pipelines, infrastructure declarations, metrics dashboards and capability-test scaffolding. Subsequent product requests—analytics storage, partner callbacks, new event types—are accommodated by small, localised edits rather than by the recreation of entire deployment artefacts. The cognitive load remains focused on the business problem; the mechanical work of packaging, provisioning and observing has been systematically removed from the critical path.

The same pattern scales. Whether the service handles a single partner’s CSV feed or participates in a multi-stage order-reconciliation flow involving a dozen intermediate services, the surrounding machinery remains identical. Consistency of structure, generation of repetitive artefacts and enforcement of hygiene through continuous integration together produce an estate that can grow without a proportional increase in operational friction. That is the practical payoff of the investment in a centralised toolkit.
Looking ahead, the same principles that make GoKit effective inside Just Eat are portable to any organisation that maintains a large collection of similarly structured services. The precise implementation will differ—different cloud providers, different event buses, different continuous-integration systems—but the underlying ideas remain constant: centralise the definition of service shape, generate everything that can be generated, enforce consistency automatically, and keep the engineer’s attention on the business problem. When those ideas are applied with discipline, the cost of creating and operating microservices falls dramatically, and the organisation’s capacity to respond to new partner or product requirements rises correspondingly.
In the end the value of a toolkit such as GoKit is measured less by the number of lines it generates than by the number of decisions it removes from the critical path of feature delivery. Every database, every topic subscription, every continuous-integration step that an engineer no longer has to configure by hand is a unit of attention that can be redirected toward the unique requirements of a partner or a product. When that redirection is systematic and reliable, the organisation as a whole becomes more responsive, more consistent and more capable of sustaining growth without a proportional increase in operational complexity.
The pizza-service narrative also illustrates a secondary benefit that is easy to overlook: the progressive enrichment of operational visibility. Because metrics, dashboards and capability tests are generated rather than hand-crafted, every new service automatically participates in the organisation’s observability and quality regimes. There is no opportunity for a service to be “forgotten” or to ship without the standard instrumentation. Consistency of tooling produces consistency of operational posture, which in turn reduces the cognitive load on both developers and on-call engineers.
Taken together, the practices embodied in GoKit demonstrate that the apparent tension between rapid delivery and long-term maintainability can be resolved by investing in the right abstractions at the platform level. When the platform absorbs the repetitive work of scaffolding, provisioning, testing and observing, individual service teams are free to move quickly without accumulating the structural debt that eventually slows every organisation that scales through pure copy-and-paste. The result is an engineering culture that can sustain high throughput while preserving the coherence and reliability required for a production system that processes hundreds of millions of orders.
The same principles that enabled Jet Connect to move from a maintenance-heavy template model to a coherent, generated estate are available to any organisation facing a similar proliferation of services. The precise technology choices—Terraform versus another infrastructure-as-code system, Kafka versus another event bus, OpenAPI versus another interface description—matter less than the architectural decision to centralise the definition of service shape and to generate the repetitive surrounding machinery. Once that decision is made and enforced, the cost of each additional service falls and the organisation’s ability to respond to new requirements rises. In an environment that processes hundreds of millions of orders, that difference is decisive.
In practical terms the toolkit converts what would otherwise be a multi-day or multi-week effort—creating a repository, wiring continuous integration, declaring infrastructure, adding metrics, writing documentation—into a sequence of minutes. The engineer’s attention remains on the unique aspects of the partner integration or the product feature. Everything else is supplied by generation and enforced by pipeline. That separation of concerns is the essential achievement, and it is the reason the platform can continue to scale without a corresponding explosion in operational overhead.
The experience of Jet Connect demonstrates that the investment in a carefully designed internal platform pays continuing dividends. Each new service inherits the accumulated learning of the entire estate; each improvement to the toolkit propagates automatically; each engineer inherits a consistent, well-instrumented starting point. The result is not merely faster delivery of individual features but a durable increase in the organisation’s capacity to absorb complexity without sacrificing reliability or developer effectiveness.
In short, GoKit is less a code generator than a deliberate architectural intervention that realigns incentives and removes friction from the path of delivery.

Links:

PostHeaderIcon [GopherConUK2025] CPU Quota Semantics and Runtime Scheduler Behavior in Containerized Environments

Lecturer

Bill Kennedy is a software engineer, technical trainer, and Managing Partner at Ardan Labs. He has authored multiple technical books on Go programming and serves as a core organizer for developer communities worldwide. His professional work focuses on high-performance backend development, system design, and training software engineering teams on runtime internals and concurrent programming semantics.

Abstract

Deploying managed language runtimes into containerized orchestration frameworks requires a comprehensive understanding of how compute limits interact with application-level scheduling primitives. This article examines the behavior of the Go runtime scheduler when executed under Kubernetes CPU limits. By analyzing thread management, operating system context switching, and the mechanics of Completely Fair Scheduler (CFS) quota enforcement, this study highlights performance degradation scenarios caused by misalignment between thread allocation and container CPU constraints. Furthermore, empirically derived benchmarking demonstrates how adjusting runtime concurrency configurations mitigates kernel-level throttling and improves request throughput in CPU-bound and IO-bound application workloads.

Micro-Architecture, Concurrency, and Context Switching Mechanics

Modern multi-core processors execute operations via clock cycles, where instruction execution frequency depends on pipelined hardware architectures. On a standard processor core running at a 3 GHz clock rate, a single nanosecond corresponds to three clock cycles. Leveraging superscalar execution pipelines, modern hardware can process up to four instructions per clock cycle on average, yielding approximately twelve instructions per nanosecond. Consequently, operational latencies—whether originating from memory access, network round trips, or kernel thread context switches—directly translate into unexecuted instruction cycles.

Operational Event Approximate Duration Lost Instruction Opportunities
OS Thread Context Switch 1,000 ns (1 µs) ~12,000 instructions
Datacenter Network Round Trip 500,000 ns (0.5 ms) ~6,000,000 instructions
Go Routine Context Switch 200 ns ~2,400 instructions

In system software, workloads are categorized as either CPU-bound or IO-bound. CPU-bound tasks execute uninterrupted mathematical or logical operations, utilizing their full operating system time slice. Under CPU-bound conditions, context switches incur overhead that degrades throughput unless application thread counts strictly match available physical cores. Conversely, IO-bound workloads frequently transition threads into blocked or waiting states due to asynchronous network calls or file interactions.

The Go runtime abstracts operating system (OS) threads through an M:N scheduler, mapping M goroutines (application-level lightweight threads) onto N OS threads managed across logical processors known as P structures. Physical CPU cores are abstracted into these P units, which hold local run queues for goroutines. The Go scheduler operates as a work-stealing system: idle P structures steal runnable goroutines from other local queues or a global run queue.

+-----------------------------------------------+
|                 OS Kernel                     |
|  [Core 0]    [Core 1]    [Core 2]    [Core 3] |
+-----------------------------------------------+
       ^          ^           ^           ^
       |          |           |           |
    [  M  ]    [  M  ]     [  M  ]     [  M  ]
       |          |           |           |
    [  P  ]    [  P  ]     [  P  ]     [  P  ]
    /     \      ...         ...         ...
 [ G ]   [ G ]

To maximize thread utilization, asynchronous system calls (such as network operations) are handled via a dedicated network poller thread. When a goroutine initiates a network read, the runtime detaches the goroutine from its current M thread and registers it with the network poller. This frees the underlying M thread to immediately execute other goroutines assigned to that logical P processor. Synchronous operations, such as blocking file system IO, force the runtime to decouple the blocking M thread from its assigned P structure and allocate or unpark a separate OS thread to keep the P processor active.

Through this abstraction, the Go runtime transforms application-level IO-bound tasks into CPU-bound operational streams from the operating system’s perspective. The OS kernel observes saturated worker threads (M), allowing them to consume allocated time slices efficiently without premature thread parking.

Kubernetes Completely Fair Scheduler (CFS) Quota Semantics

Kubernetes enforces compute resource limits using Linux control groups (cgroups) via the Completely Fair Scheduler (CFS) quota system. A CPU resource limit specified in millicores (such as 250m) translates into a time-based allocation per enforcement period. By default, the Linux kernel CFS operates on a 100-millisecond period.

Allocated Time Formula:
Allocated Time = CFS Period * (Millicores / 1000)

For an allocation of 250m across a 100 ms period, the container receives exactly 25 ms of cumulative execution time:

Allocated Time = 100 ms * (250 / 1000) = 25 ms

Crucially, the Linux CFS tracks CPU quota consumption cumulatively across all running OS threads within the container’s thread group. If an application spawns 16 threads that execute concurrently on a multi-core host system, each thread consumes physical core time simultaneously.

Quota Exhaustion Rate:
Quota Exhaustion Rate = Number of Threads * Elapsed Time

With 16 active OS threads, a 25 ms CPU quota is depleted in less than 2 ms of real time:

Exhaustion Time = 25 ms / 16 = 1.5625 ms

Once the total execution time across all threads reaches the 25 ms ceiling, the kernel CFS throttles the entire container. The container processes remain paused until the 100 ms cycle resets, resulting in severe latency spikes and degraded service throughput.

Architectural Misalignment: Go Max Procs in Container Runtimes

By default, the Go runtime initializes the number of logical processors (P) via the GOMAXPROCS variable based on system calls that query host core availability. In standard Kubernetes pod deployments without explicit runtime tuning, the runtime inspects the host node rather than container cgroup boundaries.

If a pod configured with a 250m limit is scheduled on a 16-core physical node, GOMAXPROCS defaults to 16. The runtime creates 16 logical P processors and corresponding OS threads (M).

Container Configuration: CPU Limit = 250m (25ms per 100ms cycle)
Host Infrastructure: 16 Physical Cores
Default Go Runtime Behavior: GOMAXPROCS = 16

+-------------------------------------------------------+
| 16 OS Threads (M) Executing Simultaneously           |
| [M1] [M2] [M3] [M4] [M5] [M6] ... [M16]               |
+-------------------------------------------------------+
                           |
                           v
    Consumes 25ms Quota in ~1.56ms of Real Time
                           |
                           v
+-------------------------------------------------------+
| Kernel CFS Throttles Container for Remaining ~98.4ms  |
+-------------------------------------------------------+

When incoming requests hit the container, all 16 worker threads wake up to process goroutines. The cumulative CPU time consumed by these 16 concurrent threads exhausts the 25 ms cgroup quota almost instantly. The application spends the vast majority of every 100 ms enforcement window in a kernel-throttled state.

To resolve this misalignment, the application runtime must match its logical thread capacity to its cgroup quota boundaries. Setting GOMAXPROCS=1 forces the Go scheduler to utilize a single logical P processor and one primary operating thread, executing sequential instructions over the full 25 ms window without premature multi-threaded quota depletion.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: sales-service
spec:
  template:
    spec:
      containers:
      - name: service
        image: sales-service:1.0
        env:
        - name: GOMAXPROCS
          valueFrom:
            resourceFieldRef:
              resource: limits.cpu
        resources:
          limits:
            cpu: "250m"

In Go deployments, setting GOMAXPROCS via container environment variables applies a mathematical ceiling function to convert fractional core limits into discrete thread bounds.

Experimental Evaluation and System Optimization

Empirical performance tests were conducted on a Kubernetes cluster managed via kind hosted on an 16-core machine. The microservice stack comprised an HTTP API service backed by a PostgreSQL database. Load testing was executed using automated benchmark tools transmitting synthetic HTTP workloads.

Test Configuration A: Default Core Allocation

  • CPU Limit: 250m (25 ms per 100 ms)
  • Host Cores Detected: 16
  • GOMAXPROCS: 16 (Default)

Test Configuration B: Matched Runtime Constraints

  • CPU Limit: 250m (25 ms per 100 ms)
  • Host Cores Detected: 16
  • GOMAXPROCS: 1 (Explicitly configured)

Measured Experimental Results

Metric Config A (GOMAXPROCS=16) Config B (GOMAXPROCS=1) Performance Impact
Throughput ~126 req/sec ~2,746 req/sec ~21.7x Increase
P99 Latency ~200 ms ~3.6 ms ~98.2% Reduction

Constraining the runtime thread count to align with container limits produced a 21-fold throughput increase while eliminating excessive tail latency caused by kernel CFS throttling.

Cascading Latency and Upstream Service Dependencies

In distributed microservice topologies, runtime throttling can cascade across service boundaries. During secondary experimentation, an authentication service dependency (auth-service) was assigned a restricted CPU limit (100m).

Even when the primary edge service (sales-service) was provisioned with unrestricted CPU allocations, overall request throughput dropped to baseline throttled levels. Blocking latencies introduced by the throttled upstream dependency bottlenecked the unconstrained downstream service. Diagnosing performance anomalies requires evaluating total system execution graphs rather than isolating individual application metrics.

Links:

PostHeaderIcon [KCDUK2024] Kubernetes Privilege Escalation Tactics: Unveiling Vulnerabilities and Fortifying Defenses

Iain Smart and Andrew Martin presented a compelling exploration of Kubernetes privilege escalation tactics, offering a deep dive into how both trusted and unprivileged users can exploit vulnerabilities within the system. Their discussion, a highlight of KCDUK2024, provided invaluable insights for SREs, security teams, and pentesters aiming to enhance cluster security.

The core of their presentation focused on the critical need to understand potential attack vectors and implement robust defense mechanisms. They articulated that while penetration testing Kubernetes should inherently be challenging, certain oversights can inadvertently simplify the process for malicious actors. The speakers emphasized the multifaceted nature of threats, ranging from rogue SREs and disaffected platform developers to external hostile internet citizens.

Smart and Martin meticulously guided attendees through various methods to escalate privileges, achieve persistence, and potentially wreak havoc across a cluster, all while attempting to obscure any traces of activity. Their expertise illuminated the intricate interplay of Kubernetes components and how unusual interactions or component abuse can be leveraged for unauthorized access.

Understanding Kubernetes Vulnerabilities and Exploitation

The presenters underscored the importance of comprehending the array of Kubernetes vulnerabilities that security professionals must be aware of. They elaborated on specific techniques that adversaries might employ, detailing how seemingly minor misconfigurations or overlooked edge cases can become critical points of entry. The discussion extended to identifying different adversary levels, stressing that tailoring defenses according to the threat model is paramount for effective security.

Smart and Martin provided practical insights into the most cost-effective and efficient strategies for fortifying Kubernetes clusters. They advocated for a proactive approach that encompasses preventative controls, such as static analysis on deployed artifacts and meticulous enumeration of Role-Based Access Control (RBAC). The emphasis was not solely on prevention but also on robust detective controls, including monitoring external traffic through split-horizon DNS to identify suspicious outbound communications.

Mitigating Risks and Ensuring Robust Security

A key takeaway from Smart and Martin’s presentation was the critical role of remediative controls. These controls are designed to detect ongoing attacks and initiate automated responses, such as node draining and shutdown procedures, to prevent data exfiltration. Despite implementing a comprehensive suite of preventative, detective, and remediative measures, they acknowledged that complete isolation and absolute security are unattainable ideals. The ever-evolving threat landscape, exemplified by nation-state attacks and sophisticated persistent threats, necessitates continuous vigilance and adaptation.

The speakers concluded by reiterating that detection is as crucial as prevention in the complex puzzle of cloud-native security. They highlighted that while various security tools and practices are available, the dynamic nature of threats requires an ongoing commitment to learning, monitoring, and adapting. Their insights provided a valuable roadmap for organizations striving to secure their Kubernetes environments against an increasingly sophisticated array of threats.

Links:

PostHeaderIcon [DevoxxFR2025] Boosting Java Application Startup Time: JVM and Framework Optimizations

In the world of modern application deployment, particularly in cloud-native and microservice architectures, fast startup time is a crucial factor impacting scalability, resilience, and cost efficiency. Slow-starting applications can delay deployments, hinder auto-scaling responsiveness, and consume resources unnecessarily. Olivier Bourgain, in his presentation, delved into strategies for significantly accelerating the startup time of Java applications, focusing on optimizations at both the Java Virtual Machine (JVM) level and within popular frameworks like Spring Boot. He explored techniques ranging from garbage collection tuning to leveraging emerging technologies like OpenJDK’s Project Leyden and Spring AOT (Ahead-of-Time Compilation) to make Java applications lighter, faster, and more efficient from the moment they start.

The Importance of Fast Startup

Olivier began by explaining why fast startup time matters in modern environments. In microservices architectures, applications are frequently started and stopped as part of scaling events, deployments, or rolling updates. A slow startup adds to the time it takes to scale up to handle increased load, potentially leading to performance degradation or service unavailability. In serverless or function-as-a-service environments, cold starts (the time it takes for an idle instance to become ready) are directly impacted by application startup time, affecting latency and user experience. Faster startup also improves developer productivity by reducing the waiting time during local development and testing cycles. Olivier emphasized that optimizing startup time is no longer just a minor optimization but a fundamental requirement for efficient cloud-native deployments.

JVM and Garbage Collection Optimizations

Optimizing the JVM configuration and understanding garbage collection behavior are foundational steps in improving Java application startup. Olivier discussed how different garbage collectors (like G1, Parallel, or ZGC) can impact startup time and memory usage. Tuning JVM arguments related to heap size, garbage collection pauses, and just-in-time (JIT) compilation tiers can influence how quickly the application becomes responsive. While JIT compilation is crucial for long-term performance, it can introduce startup overhead as the JVM analyzes and optimizes code during initial execution. Techniques like Class Data Sharing (CDS) were mentioned as a way to reduce startup time by sharing pre-processed class metadata between multiple JVM instances. Olivier provided practical tips and configurations for optimizing JVM settings specifically for faster startup, balancing it with overall application performance.

Framework Optimizations: Spring Boot and Beyond

Popular frameworks like Spring Boot, while providing immense productivity benefits, can sometimes contribute to longer startup times due to their extensive features and reliance on reflection and classpath scanning during initialization. Olivier explored strategies within the Spring ecosystem and other frameworks to mitigate this. He highlighted Spring AOT (Ahead-of-Time Compilation) as a transformative technology that analyzes the application at build time and generates optimized code and configuration, reducing the work the JVM needs to do at runtime. This can significantly decrease startup time and memory footprint, making Spring Boot applications more suitable for resource-constrained environments and serverless deployments. Project Leyden in OpenJDK, aiming to enable static images and further AOT compilation for Java, was also discussed as a future direction for improving startup performance at the language level. Olivier demonstrated how applying these framework-specific optimizations and leveraging AOT compilation can have a dramatic impact on the startup speed of Java applications, making them competitive with applications written in languages traditionally known for faster startup.

Links:

PostHeaderIcon [KCDUK2024] The Operator Antipattern | Gerald Schmidt

At KCDUK2024, Gerald Schmidt, data platform engineering lead at DS Smith, delivered a candid reflection on the Kubernetes operator pattern, questioning its efficacy in modern platform engineering. Drawing from his experience at Babylon Health, Go City, and Thoughtworks, Gerald shared a journey from enthusiasm to disillusionment, highlighting how operators, despite their promise, often introduce complexity and fragility that outweigh their benefits.

The Allure and Pitfalls of Operators

Gerald opened with his initial excitement for operators, which promised to extend Kubernetes’ API for managing complex stateful applications. He referenced Brandon Phillips’ vision of operators embedding domain expertise to automate tasks, such as managing databases or message queues. However, Gerald’s experience revealed that operators rarely delivered on this promise. Instead, they introduced tight coupling between custom resource definitions (CRDs) and controllers, creating operational challenges.

A pivotal anecdote from his time at a health startup illustrated this. Gerald’s team built operators for data flows, only to face disaster when CRDs were inadvertently deleted, requiring a weekend-long recovery. This incident led to their team losing deployment privileges, underscoring the risks of operators in environments with shrinking teams and heightened scrutiny. Gerald argued that operators, while powerful, often fail to match the robustness of managed services like AWS RDS, which offer superior durability and availability.

Persistent Volumes Over Stateful Applications

Gerald challenged the notion that Kubernetes struggles with stateful applications, suggesting the real issue lies with persistent volumes. Solutions like Thanos and WarpStream, which leverage object storage, address this effectively without relying on operators. He highlighted the maturity of the object storage ecosystem, with providers like AWS S3, Google Cloud Storage, and Wasabi offering cost-effective, reliable alternatives. The Container Object Storage Interface (COSI), though still emerging, promises further simplification, reducing the need for complex operator-based solutions.

He contrasted this with the operator pattern’s limitations, particularly around CRDs. Unlike controllers, which integrate seamlessly with Kubernetes’ control loop, CRDs require administrative privileges and complicate versioning and upgrades. Gerald cited the AWS Controllers for Kubernetes (ACK), still in v1alpha1, as an example of versioning challenges, where upgrades introduce new failure modes and complexity.

Toward Simpler, Developer-Friendly Solutions

Reflecting on developer experience, Gerald advocated for alternatives like config maps over CRDs. He praised Grafana’s approach, which uses config maps for dashboards, avoiding the need for administrative access and simplifying deployment. Similarly, older Prometheus setups allowed developers to control scraping with simple annotations, granting autonomy without the overhead of CRDs. Gerald argued that controllers alone, paired with config maps, provide most of the operator’s benefits without the downsides.

He concluded with a decision framework: only pursue the operator pattern if the use case justifies a domain-specific language. For operational tasks with low blast radius, like Kyverno’s policy enforcement, operators excel. However, for development teams, Gerald recommended avoiding custom operators, as they often lead to wasted effort and fragility, as seen in his team’s abandoned six-month project.

Links:

PostHeaderIcon [DevoxxGR2025] Email Re-Platforming Case Study

George Gkogkolis from Travelite Group shared a 15-minute case study at Devoxx Greece 2025 on re-platforming to process 1 million emails per hour.

The Challenge

Travelite Group, a global OTA handling flight tickets in 75 countries, processes 350,000 emails daily, expected to hit 2 million. Previously, a SaaS ticketing system struggled with growing traffic, poor licensing, and subpar user experience. Sharding the system led to complex agent logins and multiplexing issues with the booking engine. Market research revealed no viable alternatives, as vendors’ licensing models couldn’t handle the scale, prompting an in-house solution.

The New Platform

The team built a cloud-native, microservices-based platform within a year, going live in December 2024. It features a receiving app, a React-based web UI with Mantine Dev, a Spring Boot backend, and Amazon DocumentDB, integrated with Amazon SES and S3. Emails land in a Postfix server, are stored in S3, and processed via EventBridge and SQS. Data migration was critical, moving terabytes of EML files and databases in under two months, achieving a peak throughput of 1 million emails per hour by scaling to 50 receiver instances.

Lessons Learned

Starting with migration would have eased performance optimization, as synthetic data didn’t match production scale. Cloud-native deployment simplified scaling, and a backward-compatible API eased integration. Open standards (EML, Open API) ensured reliability. Future plans include AI and LLM enhancements by 2025, automating domain allocation for scalability.

Links