Recent Posts
Archives

Posts Tagged ‘devops’

PostHeaderIcon [MunchenJUG] Reliability in Enterprise Software: A Critical Analysis of Automated Testing in Spring Boot Ecosystems (27/Oct/2025)

Lecturer

Philip Riecks is an independent software consultant and educator specializing in Java, Spring Boot, and cloud-native architectures. With over seven years of professional experience in the software industry, Philip has established himself as a prominent voice in the Java ecosystem through his platform, Testing Java Applications Made Simple. He is a co-author of the influential technical book Stratospheric: From Zero to Production with Spring Boot and AWS, which bridges the gap between local development and production-ready cloud deployments. In addition to his consulting work, he produces extensive educational content via his blog and YouTube channel, focusing on demystifying complex testing patterns for enterprise developers.

Abstract

In the contemporary landscape of rapid software delivery, automated testing serves as the primary safeguard for application reliability and maintainability. This article explores the methodologies for demystifying testing within the Spring Boot framework, moving beyond superficial unit tests toward a comprehensive strategy that encompasses integration and slice testing. By analyzing the “Developer’s Dilemma”—the friction between speed of delivery and the confidence provided by a robust test suite—this analysis identifies key innovations such as the “Testing Pyramid” and specialized Spring Boot test slices. The discussion further examines the technical implications of external dependency management through tools like Testcontainers and WireMock, advocating for a holistic approach that treats test code with the same rigor as production logic.

The Paradigm Shift in Testing Methodology

Traditional software development often relegated testing to a secondary phase, frequently outsourced to separate quality assurance departments. However, the rise of DevOps and continuous integration has necessitated a shift toward “test-driven” or “test-enabled” development. Philip Riecks identifies that the primary challenge for developers is not the lack of tools, but the lack of a clear strategy. Testing is often perceived as a bottleneck rather than an accelerator.

The methodology proposed focuses on the Testing Pyramid, which prioritizes a high volume of fast, isolated unit tests at the base, followed by a smaller number of integration tests, and a minimal set of end-to-end (E2E) tests at the apex. The innovation in Spring Boot testing lies in its ability to provide “Slice Testing,” allowing developers to load only specific parts of the application context (e.g., the web layer or the data access layer) rather than the entire infrastructure. This approach significantly reduces test execution time while maintaining high fidelity.

Architectural Slicing and Context Management

One of the most powerful features of the Spring Boot ecosystem is its refined support for slice testing via annotations. This allows for an analytical approach to testing where the scope of the test is strictly defined by the architectural layer under scrutiny.

  1. Web Layer Testing: Using @WebMvcTest, developers can test REST controllers without launching a full HTTP server. This slice provides a mocked environment where the web infrastructure is active, but business services are replaced by mocks (e.g., using @MockBean).
  2. Data Access Testing: The @DataJpaTest annotation provides a specialized environment for testing JPA repositories. It typically uses an in-memory database by default, ensuring that database interactions are verified without the overhead of a production-grade database.
  3. JSON Serialization: @JsonTest isolates the serialization and deserialization logic, ensuring that data structures correctly map to their JSON representations.

This granular control prevents “Context Bloat,” where tests become slow and brittle due to the unnecessary loading of the entire application environment.

Code Sample: A Specialized Controller Test Slice

@WebMvcTest(UserRegistrationController.class)
class UserRegistrationControllerTest {

    @Autowired
    private MockMvc mockMvc;

    @MockBean
    private UserRegistrationService registrationService;

    @Test
    void shouldRegisterUserSuccessfully() throws Exception {
        mockMvc.perform(post("/api/users")
                .contentType(MediaType.APPLICATION_JSON)
                .content("{\"username\": \"priecks\", \"email\": \"philip@example.com\"}"))
                .andExpect(status().isCreated());
    }
}

Managing External Dependencies: Testcontainers and WireMock

A significant hurdle in integration testing is the reliance on external systems such as databases, message brokers, or third-party APIs. Philip emphasizes the move away from “In-Memory” databases (like H2) for testing production-grade applications, citing the risk of “Environment Parity” issues where H2 behaves differently than a production PostgreSQL instance.

The integration of Testcontainers allows developers to spin up actual Docker instances of their production infrastructure during the test lifecycle. This ensures that the code is tested against the exact same database engine used in production. Similarly, WireMock is utilized to simulate external HTTP APIs, allowing for the verification of fault-tolerance mechanisms like retries and circuit breakers without depending on the availability of the actual external service.

Consequences of Testing on Long-term Maintainability

The implications of a robust testing strategy extend far beyond immediate bug detection. A well-tested codebase enables fearless refactoring. When developers have a “safety net” of automated tests, they can update dependencies, optimize algorithms, or redesign components with the confidence that existing functionality remains intact.

Furthermore, Philip argues that the responsibility for quality must lie with the engineer who writes the code. In an “On-Call” culture, the developer who builds the system also runs it. This ownership model, supported by automated testing, transforms software engineering from a process of “handing over” code to one of “carefully crafting” resilient systems.

Conclusion

Demystifying Spring Boot testing requires a transition from viewing tests as a chore to seeing them as a fundamental engineering discipline. By leveraging architectural slices, managing dependencies with Testcontainers, and adhering to the Testing Pyramid, developers can build applications that are not only functional but also sustainable. The ultimate goal is to reach a state where testing provides joy through the confidence it instills, ensuring that the software remains a robust asset for the enterprise rather than a source of technical debt.

Links:

PostHeaderIcon [GopherConUK2025] Go Module Hygiene: Keeping go.sum and go.mod in Check

Lecturer

Emily Achieng is a software engineer specialising in DevOps. Originally from Kenya and based in Switzerland, she has previously spoken at the inaugural GopherCon South Africa. Her professional experience covers server-side development and cloud-native technologies, including Kubernetes. She draws on concrete operational experience with dependency management in production systems and shares practical techniques that keep Go projects maintainable and secure over time.

Abstract

This article examines the practical discipline of Go module hygiene. It treats go.mod and go.sum as the foundational artefacts that underwrite reproducible builds, dependency consistency and supply-chain security. Drawing on common symptoms of neglect—slow builds, inflated binaries, mysterious runtime errors and version conflicts—the discussion presents concrete remediation techniques, security practices and sustainable habits. Emphasis is placed on regular pruning, informed selection of third-party packages, automated scanning and the progressive embedding of hygiene into continuous-integration pipelines so that healthy modules become the default rather than the exception.

The Dynamic Duo and the Cost of Neglect

Every Go project rests on two files that rarely receive the attention they deserve. The go.mod file functions as a meticulous product manager: it records the module path, the required language version and the precise set of direct and indirect dependencies that the project claims. The go.sum file acts as a vigilant security guard: it stores cryptographic checksums that guarantee each downloaded module matches the version originally resolved by the toolchain. Together they constitute the bedrock of reproducible builds and the first line of defence against supply-chain compromise.

Yet even the most diligent files can become overwhelmed. Over time a go.mod accumulates entries that are no longer referenced, outdated versions that harbour known vulnerabilities, and transitive dependencies that pull in conflicting requirements. The resulting clutter manifests in several observable symptoms. Builds that once completed in seconds stretch into coffee-break or even lunch-break durations while the toolchain sifts through unnecessary packages and resolves version constraints. Binary sizes inflate because large libraries are pulled in for a handful of functions, increasing deployment weight and memory pressure. Runtime panics appear in apparently unrelated code paths when two versions of the same package fight for dominance at link or run time. Version conflicts leave the toolchain forced to choose among incompatible requirements, producing cryptic errors that are difficult to diagnose and that surface only under particular load or configuration conditions.

These symptoms are not merely inconveniences. Outdated dependencies leave open windows through which known vulnerabilities can be exploited. Bloated dependency graphs enlarge the attack surface and complicate auditing. The cumulative effect is a form of technical debt that is specifically dependency debt—real, measurable and costly to repay once it has accumulated. Recognising the symptoms early is therefore the first practical skill of module hygiene.

Recognising Symptoms and Performing First-Line Cleanup

The earliest warning signs are often temporal and spatial. When a build that should be instantaneous requires an extended pause, or when a modest service suddenly demands additional storage for its binary, the module graph is almost certainly carrying dead weight. Mysterious errors that surface after an apparently unrelated change frequently point to transitive conflicts. Missing-package messages that appear even though the parent module is present indicate incomplete or inconsistent resolution. Dependency conflicts in which different parts of the project demand incompatible versions of the same module force the toolchain into difficult choices that can produce runtime surprises.

The first and most immediate remedy is the command go mod tidy. Analogous to a spring-cleaning exercise that sorts clothing into keep and donate piles, it removes unused requirements, adds any that are actually imported by the current source, and rewrites both go.mod and go.sum into a minimal consistent state. The effect is usually immediate: binary size shrinks, build times improve and the module graph becomes legible again. Because the command is fast and non-destructive, it can be run frequently without ceremony.

Complementary inspection tools deepen the diagnosis. go mod graph renders the full dependency tree, making visible the social network of packages and revealing which indirect modules have entered through which direct ones. The resulting graph is invaluable when an unexpected package appears or when a vulnerability is reported in a transitive dependency. go mod why answers the precise question of why a particular module is present, tracing the import path that pulled it in and thereby empowering informed decisions rather than blind removal. go list -m all supplies a complete inventory of every module, direct and transitive, functioning as a magnifying glass over the entire dependency set.

When absolute stability is required—particularly inside continuous-integration pipelines—version pinning becomes appropriate. By fixing a dependency to an exact version rather than a range, the build is insulated from unexpected upstream changes. The trade-off is that updates must be deliberate; the benefit is the elimination of “it worked yesterday / it worked on my machine” surprises. In multi-environment scenarios, replace directives can point a dependency at a local or forked copy without permanently altering the published module graph, provided the engineer remains mindful of the temporary nature of the redirection.

Selecting Dependencies and Guarding the Supply Chain

Before any third-party package is added, two questions should be asked. First, does the standard library already solve the problem? Embracing the standard library reduces external surface area, improves performance and eliminates an entire class of maintenance burden. Many common tasks—HTTP clients, JSON handling, cryptography, compression—are already well covered; reaching for an external package should be a conscious choice rather than a default reflex. Second, if an external package is genuinely required, is it actively maintained? Indicators include recent commits, responsive issue tracking, clear documentation, a healthy community of contributors and evidence that critical bugs are addressed promptly. Packages last updated years ago with unresolved critical issues leave the consumer effectively alone when problems arise.

Security considerations reinforce the same discipline. Each dependency is a door into the application; an outdated or compromised package is an unlocked door. Supply-chain attacks demonstrate that even widely trusted packages can be subverted. The go.sum checksums provide cryptographic verification that the downloaded code matches the expected content, functioning as an identity check at the door. Without them, an attacker who could place a malicious module under the same path would succeed unnoticed. Automated vulnerability scanners integrated into the continuous-integration pipeline act as tireless security patrols that never sleep and catch problems early, before they reach production.

Vendoring (go mod vendor) creates a local snapshot of the dependency tree. While it transfers the responsibility for updates onto the project itself, it grants absolute control over the exact code that will be compiled, removing reliance on external registries at build time. The analogy is stocking a pantry rather than making a trip to the grocery for every meal: convenience and control are gained at the price of periodic restocking.

Sustainable Habits and Automation

Hygiene is not a heroic one-time act; it is a set of small, repeated practices. Scheduling regular reviews—perhaps aligned with sprint cycles or release trains—prevents the accumulation of dead weight. Updating one dependency at a time is far less terrifying than attempting a wholesale refresh of an entire graph. Documenting the rationale for each non-obvious choice aids future maintainers and accelerates onboarding of new team members. Comprehensive tests, run before any version bump is merged, protect against regressions that would otherwise surface only in production.

The ultimate goal is to make the healthy state the default state. Continuous-integration pipelines can fail a build when go mod tidy would produce a diff, thereby enforcing cleanliness without manual effort. Vulnerability scanners can run on every pull request and block merges that introduce known issues. Dependency-update bots can open carefully scoped pull requests that are reviewed and tested like any other change. When these mechanisms are in place, module hygiene becomes background infrastructure rather than foreground labour, freeing engineers to focus on product work while the toolchain quietly maintains the integrity of the dependency graph.

In summary, go.mod and go.sum are small files with large consequences. Treating them with the same care given to application code—spotting symptoms early, pruning regularly, choosing dependencies deliberately, verifying integrity and automating the routine—keeps projects fast, lean, reproducible and secure as they grow. Dependency debt is real; the habits that prevent it are correspondingly valuable.
The practice of module hygiene ultimately rests on a shift in mindset. Dependencies are not free; each one carries a maintenance and security cost that compounds over the lifetime of a project. Treating the addition of a new module as a deliberate architectural decision rather than a casual convenience changes the calculus. When the standard library can serve, it should. When an external package is required, its maintenance status, community health and security track record become first-class evaluation criteria. When the graph inevitably grows, regular pruning and automated enforcement keep it from becoming unmanageable.

Automation does not replace judgement; it amplifies it. A pipeline that fails on an untidy go.mod forces the conversation about necessity to happen at the moment of change rather than months later when the mess has become entrenched. A scanner that surfaces a vulnerability in a transitive dependency gives the team the information needed to decide whether to upgrade, replace or accept the risk. The combination of human review and machine enforcement produces a sustainable equilibrium that neither pure manual diligence nor pure automation can achieve alone.
In practice the most effective teams treat module hygiene as a continuous background process rather than a periodic cleanup project. They embed tidy checks into every pull-request pipeline, maintain a short list of approved packages for common tasks, and schedule lightweight dependency reviews alongside ordinary sprint work. The result is that the module graph remains close to minimal at all times, security advisories can be acted upon promptly, and new team members inherit a codebase whose dependency story is still intelligible. The alternative—allowing the graph to grow unchecked until a crisis forces a heroic clean-up—is both more expensive and more risky.
The long-term health of a Go codebase is inseparable from the health of its module graph. Teams that invest in clear ownership of go.mod and go.sum, that treat every new dependency as a conscious decision, and that automate the enforcement of cleanliness discover that many of the classic pains of dependency management simply cease to appear. Builds remain fast, binaries remain lean, security reviews remain tractable, and the mental model of the system stays within the grasp of the engineers who must maintain it. That outcome is not accidental; it is the product of deliberate, sustained hygiene.
Practical experience repeatedly confirms that the cost of prevention is lower than the cost of remediation. A few minutes spent running go mod tidy, inspecting the graph, or verifying the maintenance status of a candidate package routinely save hours of later debugging and security response. When those minutes are institutionalised through pipeline checks and team norms, the savings compound across every service and every release. The discipline is therefore not an optional polish; it is a core engineering practice that directly supports velocity, reliability and security.
Ultimately the health of the module graph is a leading indicator of the health of the codebase itself. Teams that keep their dependencies lean, current and well-understood tend also to keep their application code modular, their tests meaningful and their release processes predictable. The reverse is equally true: a neglected go.mod is often the first visible sign of deeper accumulation of technical debt. By treating module hygiene as a first-class concern, organisations protect not only their supply chain but the long-term evolvability of the systems they build.
The path from a cluttered, slow and fragile module graph to a clean, fast and trustworthy one is incremental. It begins with recognition of the symptoms, proceeds through the disciplined use of the available tooling, and is sustained by the institutionalisation of small, regular practices. Organisations that follow this path discover that the investment repays itself many times over in reduced incident response, faster onboarding and greater confidence in every release.

Links:

PostHeaderIcon [NDCOslo2024] Hub-Spoke Virtual Networks in Azure – Bastiaan Wassenaar

In the labyrinthine landscape of Azure’s azure architecture, where connectivity contends with compliance, Bastiaan Wassenaar, a cloud custodian and connectivity connoisseur, clarifies the conundrum of hub-spoke virtual networks. As a Dutch devops dynamo, Bastiaan blueprints the bedrock—endpoints, peering, policies—propelling practitioners past perplexities in private provisioning. His session, a symphony of safeguards and setups, spotlights the saga from service sentinels to spoke sanctuaries, ensuring egress elegance and ingress integrity.

Bastiaan banters on boredom’s behalf: vnet vexations vex veterans, yet victory vaults with vigilance. He heralds history: vnet’s 2014 genesis, endpoints’ evolution, private peers’ precision—pivoting from public perils to partitioned paradises.

Foundations of Fortification: Endpoints and Evolutions

Service endpoints erect ramparts: subnet sentries shielding storage, SQL sans sprawl. Bastiaan bewails bandwidth burdens—10Gbps ceilings—yet blesses them as basics, bridging to private endpoints’ purity: dedicated daisy-chains, DNS delegations demystified.

Hub-spoke’s heartbeat: central hub harboring firewalls, spokes siphoning spokes—peering propagates prefixes, UDRs usher unicast. Bastiaan blueprints: Azure Firewall’s fabric, forced tunneling fortifying flows.

Orchestrating the Orbit: Peering, Policies, and Proxies

Peering’s pact: global gateways, transitive taboos—spokes supplicate hubs for harmony. Bastiaan bemoans BGP’s burdens—bidirectional broadcasts—yet bows to basics: static routes suffice for simplicity.

Policies propel protection: NSGs nestle at NICs, FW’s finesse filters flows. Bastiaan broadcasts best bets: hub’s hegemony, spokes’ seclusion—egress egressing exclusively, ingress inspecting intently.

DNS’s Dominion: Delegations and Dilemmas

DNS dances delicately: private endpoints’ FQDNs, hub’s handlers hijacking queries. Bastiaan bemoans blunders—external IPs eclipsing internals—yet extols overrides: custom configurations, conditional forwarding.

His hack: hosts’ hacks for haste, yet hub’s hegemony harmonizes hordes. Bastiaan broadcasts: reboot realms for resolution—vnet’s vicissitudes vanquished.

Victory’s Vista: Vigilance in Vastness

Bastiaan’s benediction: hub-spoke as haven, harmonizing hazards—history heeded, hurdles hurdled. His hurrah: harness helpers, heed heuristics—Azure’s arsenal awaits.

Links:

PostHeaderIcon [GopherConUK2025] CPU Quota Semantics and Runtime Scheduler Behavior in Containerized Environments

Lecturer

Bill Kennedy is a software engineer, technical trainer, and Managing Partner at Ardan Labs. He has authored multiple technical books on Go programming and serves as a core organizer for developer communities worldwide. His professional work focuses on high-performance backend development, system design, and training software engineering teams on runtime internals and concurrent programming semantics.

Abstract

Deploying managed language runtimes into containerized orchestration frameworks requires a comprehensive understanding of how compute limits interact with application-level scheduling primitives. This article examines the behavior of the Go runtime scheduler when executed under Kubernetes CPU limits. By analyzing thread management, operating system context switching, and the mechanics of Completely Fair Scheduler (CFS) quota enforcement, this study highlights performance degradation scenarios caused by misalignment between thread allocation and container CPU constraints. Furthermore, empirically derived benchmarking demonstrates how adjusting runtime concurrency configurations mitigates kernel-level throttling and improves request throughput in CPU-bound and IO-bound application workloads.

Micro-Architecture, Concurrency, and Context Switching Mechanics

Modern multi-core processors execute operations via clock cycles, where instruction execution frequency depends on pipelined hardware architectures. On a standard processor core running at a 3 GHz clock rate, a single nanosecond corresponds to three clock cycles. Leveraging superscalar execution pipelines, modern hardware can process up to four instructions per clock cycle on average, yielding approximately twelve instructions per nanosecond. Consequently, operational latencies—whether originating from memory access, network round trips, or kernel thread context switches—directly translate into unexecuted instruction cycles.

Operational Event Approximate Duration Lost Instruction Opportunities
OS Thread Context Switch 1,000 ns (1 µs) ~12,000 instructions
Datacenter Network Round Trip 500,000 ns (0.5 ms) ~6,000,000 instructions
Go Routine Context Switch 200 ns ~2,400 instructions

In system software, workloads are categorized as either CPU-bound or IO-bound. CPU-bound tasks execute uninterrupted mathematical or logical operations, utilizing their full operating system time slice. Under CPU-bound conditions, context switches incur overhead that degrades throughput unless application thread counts strictly match available physical cores. Conversely, IO-bound workloads frequently transition threads into blocked or waiting states due to asynchronous network calls or file interactions.

The Go runtime abstracts operating system (OS) threads through an M:N scheduler, mapping M goroutines (application-level lightweight threads) onto N OS threads managed across logical processors known as P structures. Physical CPU cores are abstracted into these P units, which hold local run queues for goroutines. The Go scheduler operates as a work-stealing system: idle P structures steal runnable goroutines from other local queues or a global run queue.

+-----------------------------------------------+
|                 OS Kernel                     |
|  [Core 0]    [Core 1]    [Core 2]    [Core 3] |
+-----------------------------------------------+
       ^          ^           ^           ^
       |          |           |           |
    [  M  ]    [  M  ]     [  M  ]     [  M  ]
       |          |           |           |
    [  P  ]    [  P  ]     [  P  ]     [  P  ]
    /     \      ...         ...         ...
 [ G ]   [ G ]

To maximize thread utilization, asynchronous system calls (such as network operations) are handled via a dedicated network poller thread. When a goroutine initiates a network read, the runtime detaches the goroutine from its current M thread and registers it with the network poller. This frees the underlying M thread to immediately execute other goroutines assigned to that logical P processor. Synchronous operations, such as blocking file system IO, force the runtime to decouple the blocking M thread from its assigned P structure and allocate or unpark a separate OS thread to keep the P processor active.

Through this abstraction, the Go runtime transforms application-level IO-bound tasks into CPU-bound operational streams from the operating system’s perspective. The OS kernel observes saturated worker threads (M), allowing them to consume allocated time slices efficiently without premature thread parking.

Kubernetes Completely Fair Scheduler (CFS) Quota Semantics

Kubernetes enforces compute resource limits using Linux control groups (cgroups) via the Completely Fair Scheduler (CFS) quota system. A CPU resource limit specified in millicores (such as 250m) translates into a time-based allocation per enforcement period. By default, the Linux kernel CFS operates on a 100-millisecond period.

Allocated Time Formula:
Allocated Time = CFS Period * (Millicores / 1000)

For an allocation of 250m across a 100 ms period, the container receives exactly 25 ms of cumulative execution time:

Allocated Time = 100 ms * (250 / 1000) = 25 ms

Crucially, the Linux CFS tracks CPU quota consumption cumulatively across all running OS threads within the container’s thread group. If an application spawns 16 threads that execute concurrently on a multi-core host system, each thread consumes physical core time simultaneously.

Quota Exhaustion Rate:
Quota Exhaustion Rate = Number of Threads * Elapsed Time

With 16 active OS threads, a 25 ms CPU quota is depleted in less than 2 ms of real time:

Exhaustion Time = 25 ms / 16 = 1.5625 ms

Once the total execution time across all threads reaches the 25 ms ceiling, the kernel CFS throttles the entire container. The container processes remain paused until the 100 ms cycle resets, resulting in severe latency spikes and degraded service throughput.

Architectural Misalignment: Go Max Procs in Container Runtimes

By default, the Go runtime initializes the number of logical processors (P) via the GOMAXPROCS variable based on system calls that query host core availability. In standard Kubernetes pod deployments without explicit runtime tuning, the runtime inspects the host node rather than container cgroup boundaries.

If a pod configured with a 250m limit is scheduled on a 16-core physical node, GOMAXPROCS defaults to 16. The runtime creates 16 logical P processors and corresponding OS threads (M).

Container Configuration: CPU Limit = 250m (25ms per 100ms cycle)
Host Infrastructure: 16 Physical Cores
Default Go Runtime Behavior: GOMAXPROCS = 16

+-------------------------------------------------------+
| 16 OS Threads (M) Executing Simultaneously           |
| [M1] [M2] [M3] [M4] [M5] [M6] ... [M16]               |
+-------------------------------------------------------+
                           |
                           v
    Consumes 25ms Quota in ~1.56ms of Real Time
                           |
                           v
+-------------------------------------------------------+
| Kernel CFS Throttles Container for Remaining ~98.4ms  |
+-------------------------------------------------------+

When incoming requests hit the container, all 16 worker threads wake up to process goroutines. The cumulative CPU time consumed by these 16 concurrent threads exhausts the 25 ms cgroup quota almost instantly. The application spends the vast majority of every 100 ms enforcement window in a kernel-throttled state.

To resolve this misalignment, the application runtime must match its logical thread capacity to its cgroup quota boundaries. Setting GOMAXPROCS=1 forces the Go scheduler to utilize a single logical P processor and one primary operating thread, executing sequential instructions over the full 25 ms window without premature multi-threaded quota depletion.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: sales-service
spec:
  template:
    spec:
      containers:
      - name: service
        image: sales-service:1.0
        env:
        - name: GOMAXPROCS
          valueFrom:
            resourceFieldRef:
              resource: limits.cpu
        resources:
          limits:
            cpu: "250m"

In Go deployments, setting GOMAXPROCS via container environment variables applies a mathematical ceiling function to convert fractional core limits into discrete thread bounds.

Experimental Evaluation and System Optimization

Empirical performance tests were conducted on a Kubernetes cluster managed via kind hosted on an 16-core machine. The microservice stack comprised an HTTP API service backed by a PostgreSQL database. Load testing was executed using automated benchmark tools transmitting synthetic HTTP workloads.

Test Configuration A: Default Core Allocation

  • CPU Limit: 250m (25 ms per 100 ms)
  • Host Cores Detected: 16
  • GOMAXPROCS: 16 (Default)

Test Configuration B: Matched Runtime Constraints

  • CPU Limit: 250m (25 ms per 100 ms)
  • Host Cores Detected: 16
  • GOMAXPROCS: 1 (Explicitly configured)

Measured Experimental Results

Metric Config A (GOMAXPROCS=16) Config B (GOMAXPROCS=1) Performance Impact
Throughput ~126 req/sec ~2,746 req/sec ~21.7x Increase
P99 Latency ~200 ms ~3.6 ms ~98.2% Reduction

Constraining the runtime thread count to align with container limits produced a 21-fold throughput increase while eliminating excessive tail latency caused by kernel CFS throttling.

Cascading Latency and Upstream Service Dependencies

In distributed microservice topologies, runtime throttling can cascade across service boundaries. During secondary experimentation, an authentication service dependency (auth-service) was assigned a restricted CPU limit (100m).

Even when the primary edge service (sales-service) was provisioned with unrestricted CPU allocations, overall request throughput dropped to baseline throttled levels. Blocking latencies introduced by the throttled upstream dependency bottlenecked the unconstrained downstream service. Diagnosing performance anomalies requires evaluating total system execution graphs rather than isolating individual application metrics.

Links:

PostHeaderIcon [DevoxxFR2025] Alert, Everything’s Burning! Mastering Technical Incidents

In the fast-paced world of technology, technical incidents are an unavoidable reality. When systems fail, the ability to quickly detect, diagnose, and resolve issues is paramount to minimizing impact on users and the business. Alexis Chotard, Laurent Leca, and Luc Chmielowski from PayFit shared their invaluable experience and strategies for mastering technical incidents, even as a rapidly scaling “unicorn” company. Their presentation went beyond just technical troubleshooting, delving into the crucial aspects of defining and evaluating incidents, effective communication, product-focused response, building organizational resilience, managing on-call duties, and transforming crises into learning opportunities through structured post-mortems.

Defining and Responding to Incidents

The first step in mastering incidents is having a clear understanding of what constitutes an incident and its severity. Alexis, Laurent, and Luc discussed how PayFit defines and categorizes technical incidents based on their impact on users and business operations. This often involves established severity levels and clear criteria for escalation. Their approach emphasized a rapid and coordinated response involving not only technical teams but also product and communication stakeholders to ensure a holistic approach. They highlighted the importance of clear internal and external communication during an incident, keeping relevant parties informed about the status, impact, and expected resolution time. This transparency helps manage expectations and build trust during challenging situations.

Technical Resolution and Product Focus

While quick technical mitigation to restore service is the immediate priority during an incident, the PayFit team stressed the importance of a product-focused approach. This involves understanding the user impact of the incident and prioritizing resolution steps that minimize disruption for customers. They discussed strategies for effective troubleshooting, leveraging monitoring and logging tools to quickly identify the root cause. Beyond immediate fixes, they highlighted the need to address the underlying issues to prevent recurrence. This often involves implementing technical debt reduction measures or improving system resilience as a direct outcome of incident analysis. Their experience showed that a strong collaboration between engineering and product teams is essential for navigating incidents effectively and ensuring that the user experience remains a central focus.

Organizational Resilience and Learning

Mastering incidents at scale requires building both technical and organizational resilience. The presenters discussed how PayFit has evolved its on-call rotation models to ensure adequate coverage while maintaining a healthy work-life balance for engineers. They touched upon the importance of automation in detecting and mitigating incidents faster. A core tenet of their approach was the implementation of structured post-mortems (or retrospectives) after every significant incident. These post-mortems are blameless, focusing on identifying the technical and process-related factors that contributed to the incident and defining actionable steps for improvement. By transforming crises into learning opportunities, PayFit continuously strengthens its systems and processes, reducing the frequency and impact of future incidents. Their journey over 18 months demonstrated that investing in these practices is crucial for any growing organization aiming to build robust and reliable systems.

Links:

PostHeaderIcon [DotJs2024] Becoming the Multi-armed Bandit

In the intricate ballet of software stewardship, where intuition waltzes with empiricism, resides the multi-armed bandit—a probabilistic oracle guiding choices amid uncertainty. Ben Halpern, co-founder of Forem and dev.to’s visionary steward, dissected this gem at dotJS 2024. A full-stack polymath blending code with community curation, Ben recounted its infusions across his odyssey—from parody O’Reilly covers viralizing memes to mutton-busting triumphs—framing bandits as bridges between artistic whimsy and scientific rigor, aligning devs with stakeholders in pursuit of optimal paths.

Ben’s prologue evoked dev.to’s genesis: Twitter-era jests birthing a creative agora, bandit logic A/B-testing post formats for engagement zeniths. The archetype—casino levers, pulls maximizing payouts—mirrors dev dilemmas: UI variants, feature rollouts, content cadences. Exploration probes unknowns; exploitation harvests proven yields. Ben advocated epsilon-greedy: baseline exploitation (1-ε pulls best arm), exploratory ventures (ε samples alternatives), ε tuning via Thompson sampling for contextual nuance.

Practical infusions abounded. Load balancing: bandit selects origins, favoring responsive backends. Feature flags: variants vie, metrics crown victors. Smoke tests: endpoint probes, failures demote. ML pipelines: hyperparameter hunts, models ascend via validation. Ben’s dev.to saga: title A/Bs, bandit-orchestrated, surfacing resonant headlines sans bias. Organizational strata: nascent projects revel in exploration—ideation fests yielding prototypes; maturity mandates exploitation—scaling victors, pruning pretenders. This lexicon fosters accord: explorers and scalers, once at odds, synchronize via phases, preempting pivots’ friction.

Caution tempered zeal: bandits thrive on voluminous outcomes, not trivial toggles; overzealous testing paralyzes. As AI cheapens variants—code gen’s bounty—feedback scaffolds intensify, bandits as arbiters ensuring quality amid abundance. Ben’s coda: wield judiciously, blending craft’s flair with datum’s discipline for endeavors audacious yet assured.

Algorithmic Essence and Variants

Ben unpacked epsilon-greedy’s equilibrium: 90% best-arm fealty, 10% novelty nudges; Thompson’s Bayesian ballet contextualizes. UCB (Upper Confidence Bound) optimism tempers regret, ideal for sparse signals—dev.to’s post tweaks, engagement echoes guiding refinements.

Embeddings in Dev Workflows

Balancing clusters bandit-route requests; flags unleash cohorts, telemetry triumphs. ML’s parameter quests, smoke’s sentinel sweeps—all bandit-bolstered. Ben’s ethos: binary pass-fails sideline; array assays exalt, infrastructure for insight paramount.

Strategic Alignment and Prudence

Projects arc: explore’s ideation inferno yields scale’s forge. Ben bridged divides—stakeholder symposia in bandit vernacular—averting misalignment. Overreach warns: grand stakes summon science; mundane mandates art’s alacrity, future’s variant deluge demanding deft discernment.

Links:

PostHeaderIcon [DevoxxFR2025] Simplify Your Ideas’ Containerization!

For many developers and DevOps engineers, creating and managing Dockerfiles can feel like a tedious chore. Ensuring best practices, optimizing image layers, and keeping up with security standards often add friction to the containerization process. Thomas DA ROCHA from Lenra, in his presentation, introduced Dofigen as an open-source command-line tool designed to simplify this. He demonstrated how Dofigen allows users to generate optimized and secure Dockerfiles from a simple YAML or JSON description, making containerization quicker, easier, and less error-prone, even without deep Dockerfile expertise.

The Pain Points of Dockerfiles

Thomas began by highlighting the common frustrations associated with writing and maintaining Dockerfiles. These include:
– Complexity: Writing effective Dockerfiles requires understanding various instructions, their order, and how they impact caching and layer size.
– Time Consumption: Manually writing and optimizing Dockerfiles for different projects can be time-consuming.
– Security Concerns: Ensuring that images are built securely, minimizing attack surface, and adhering to security standards can be challenging without expert knowledge.
– Lack of Reproducibility: Small changes or inconsistencies in the build environment can sometimes lead to non-reproducible images.

These challenges can slow down development cycles and increase the risk of deploying insecure or inefficient containers.

Introducing Dofigen: Dockerfile Generation Simplified

Dofigen aims to abstract away the complexities of Dockerfile creation. Thomas explained that instead of writing a Dockerfile directly, users provide a simplified description of their application and its requirements in a YAML or JSON file. This description includes information such as the base image, application files, dependencies, ports, and desired security configurations. Dofigen then takes this description and automatically generates an optimized and standards-compliant Dockerfile. This approach allows developers to focus on defining their application’s needs rather than the intricacies of Dockerfile syntax and best practices. Thomas showed a live coding demo, transforming a simple application description into a functional Dockerfile using Dofigen.

Built-in Best Practices and Security Standards

A key advantage of Dofigen is its ability to embed best practices and security standards into the generated Dockerfiles automatically. Thomas highlighted that Dofigen incorporates knowledge about efficient layering, reducing image size, and minimizing the attack surface by following recommended guidelines. This means users don’t need to be experts in Dockerfile optimization or security to create robust images. The tool handles these aspects automatically based on the provided high-level description. Thomas might have demonstrated how Dofigen helps in creating multi-stage builds or incorporating user and permission best practices, which are crucial for building secure production-ready images. By simplifying the process and baking in expertise, Dofigen empowers developers to containerize their applications quickly and confidently, ensuring that the resulting images are not only functional but also optimized and secure. The open-source nature of Dofigen also allows the community to contribute to improving its capabilities and keeping up with evolving best practices and security recommendations.

Links:

PostHeaderIcon [GopherConUK2025] Conceptual XPDB Custom Resource definition

apiVersion: policy.form3.tech/v1alpha1
kind: CrossClusterPodDisruptionBudget
metadata:
name: cockroachdb-global-pdb
spec:
maxUnavailable: 1
clusters:
– aws-region-1
– gcp-region-1
– azure-region-1
selector:
matchLabels:
app: cockroachdb


XPDB enforces global disruption caps across cluster boundaries. When node drains or pod evasions occur in one cloud provider, XPDB evaluates the global state across all environments. By guaranteeing that only a single CockroachDB pod is disrupted across the entire global infrastructure at any given moment, XPDB ensures data consensus remains fully protected during routine infrastructure maintenance.

## Operator-Driven Infrastructure and Rolling Node-Pool Management

Managing multi-tenant infrastructure across multiple jurisdictions—each containing distinct development, staging, and production environments—results in a massive expansion of node pools. Updating Kubernetes worker nodes across this matrix using traditional Infrastructure-as-Code tools like Terraform creates extreme operational friction, requiring dozens of sequential pull requests and complex deployment pipelines.

To address this management overhead, Form3 built a custom Kubernetes operator called the Cluster Lifecycle Operator. The operator runs natively inside each managed cluster and abstracts raw node-pool management behind Custom Resource Definitions (CRDs).

// Conceptual snippet of CRD controller reconciliation loop
package main

import (
“context”
“fmt”
)

type ClusterSpec struct {
Version string json:"version"
}

func ReconcileNodePool(ctx context.Context, spec ClusterSpec) error {
fmt.Printf(“Reconciling node pools to version: %s\n”, spec.Version)
// Operator logic handles sequential node drains
return nil
}


Instead of modifying declarative infrastructure files for every individual node pool across every provider, engineers update a single CRD spec controlling the target cluster version. The Cluster Lifecycle Operator handles the rolling replacement of worker nodes asynchronously, adhering to defined disruption parameters and health checks. Upgrading the global fleet requires only three sequential pull requests—promoted systematically through development, staging, and production.

## Continuous Disaster Recovery via Automated Production Chaos Injection

Conventional disaster recovery (DR) practices often rely on periodic manual failover tests driven by static documentation. In rapidly changing microservice environments, compliance-focused DR exercises fail to validate real-world resilience, as system changes can render manual playbooks obsolete immediately after testing.

Form3 enforces continuous disaster recovery validation directly within staging environments through automated fault injection. A dedicated test harness deploys synthetic client applications outside the primary infrastructure perimeter. These synthetic actors continuously execute end-to-end payment workflows against simulated payment scheme interfaces at fixed intervals.

// Custom chaos injection test runner in Go
package main

import (
“context”
“log”
“time”
)

func InjectProviderOutage(ctx context.Context, targetCloud string) error {
log.Printf(“Simulating complete network partition for provider: %s”, targetCloud)
// Inject network block rules via Chaos Mesh API
time.Sleep(30 * time.Second)
return nil
}

“`

During automated test windows, the platform uses Chaos Mesh alongside custom Go-based orchestrators to introduce disruptive scenarios:

  • Severing inter-cloud network connectivity.
  • Terminating entire database nodes or message brokers.
  • Programmatically isolating an entire cloud provider for 24 hours.

If synthetic payment processing encounters errors or breaches latency thresholds, automated alerts notify on-call platform teams. Automated daily reports detail platform behavior during fault injection cycles, verifying that the loss of an entire cloud vendor produces zero client impact.

Links:

PostHeaderIcon [RivieraDev2025] Dhruv Kumar – Platform Engineering + AI: The Next-Gen DevOps

At Riviera DEV 2025, Dhruv Kumar delivered an engaging presentation on platform engineering, a discipline reshaping software delivery by addressing modern development challenges. Stepping in for Silva Devi, Dhruv, a senior product manager at CloudBees, explored how platform engineering, augmented by artificial intelligence, streamlines workflows, enhances developer productivity, and mitigates the complexities of cloud-native environments. His talk illuminated the transformative potential of internal developer platforms (IDPs) and AI-driven automation, offering a vision for a more efficient and secure software development lifecycle (SDLC).

The Challenges of Modern Software Development

Dhruv began by highlighting the evolving responsibilities of developers, who now spend only about 11% of their time coding, according to a survey by software.com. The remaining time is consumed by non-coding tasks such as testing, deployment, and managing security vulnerabilities. The shift-left movement, while intended to empower developers by integrating testing and deployment earlier in the process, often burdens them with tasks outside their core expertise. This is compounded by the transition to cloud environments, which introduces complex microservices architectures and distributed systems, creating navigation challenges and integration headaches.

Additionally, the rise of AI has accelerated software development, increasing code volume and tool proliferation, while supply chain attacks exploit these complexities, demanding constant vigilance from developers. Dhruv emphasized that these challenges—fragmented workflows, heightened security risks, and tool overload—necessitate a new approach to streamline processes and empower teams.

Platform Engineering: A Unified Approach

Platform engineering emerges as a solution to these issues, providing a cohesive framework for software delivery. Dhruv defined it as the discipline of designing toolchains and workflows that enable self-service capabilities for engineering teams in the cloud-native era. Central to this is the concept of an internal developer platform (IDP), which integrates tools and processes across the SDLC, from coding to deployment. By establishing a common SDLC model and vocabulary, platform engineering ensures that stakeholders—developers, QA, and security teams—share a unified understanding, reducing miscommunication and enhancing actionability.

Dhruv highlighted three pillars of effective platform engineering: a standardized SDLC model, secure best practices embedded in workflows, and the freedom for developers to use familiar tools. This last point, supported by a Forbes study from September 2023, underscores that happier developers, using tools they prefer, complete tasks 10% faster. By fostering collaboration and reducing context-switching, platform engineering creates an environment where developers can focus on innovation rather than operational overhead.

AI as a Catalyst for Optimization

Artificial intelligence plays a pivotal role in amplifying platform engineering’s impact. Dhruv explained that AI’s value lies not in generating code but in filtering noise and optimizing practices. By leveraging a robust SDLC data model, AI can provide actionable insights, provided it is fed high-quality data. For instance, AI-driven testing can prioritize time-intensive issues, streamline QA processes, and run only relevant tests based on code changes, reducing costs and feedback cycles. Dhruv cited examples like AI agents identifying vulnerabilities in code components or assessing risks in production ecosystems, automating fixes where appropriate.

He also introduced the Model Context Protocol (MCP), an open standard that enables applications to provide context to large language models, enhancing AI’s ability to deliver precise recommendations. From troubleshooting CI/CD pipelines to onboarding new developers, AI, when integrated with platform engineering, empowers teams to address bottlenecks and scale efficiently in a cloud-native world.

Empowering Developers and Securing the Future

Dhruv concluded by emphasizing that platform engineering, bolstered by AI, re-engages all actors in the software delivery process, from developers to leadership. By normalizing data across tools and providing metrics like DORA (DevOps Research and Assessment), IDPs offer visibility into bottlenecks and investment opportunities. This holistic approach not only secures the tech stack against supply chain attacks but also fosters a culture of productivity and developer satisfaction.

He encouraged attendees to explore CloudBees’ platform, which exemplifies these principles by breaking free from traditional platform limitations. Dhruv’s call to action urged developers to adopt platform engineering practices, leverage AI for optimization, and provide feedback to refine these evolving methodologies, ensuring a future where software delivery is both efficient and resilient.

Links:

PostHeaderIcon [DevoxxFR2025] Dagger Modules: A Swiss Army Knife for Modern CI/CD Pipelines

Continuous Integration and Continuous Delivery (CI/CD) pipelines are the backbone of modern software development, automating the process of building, testing, and deploying applications. However, as these pipelines grow in complexity, they often become difficult to maintain, debug, and port across different execution platforms, frequently relying on verbose and platform-specific YAML configurations. Jean-Christophe Sirot, in his presentation, introduced Dagger as a revolutionary approach to CI/CD, allowing pipelines to be written as code, executable locally, testable, and portable. He explored Dagger Functions and Dagger Modules as key concepts for creating and sharing reusable, language-agnostic components for CI/CD workflows, positioning Dagger as a versatile “Swiss Army knife” for modernizing these critical pipelines.

The Pain Points of Traditional CI/CD

Jean-Christophe began by outlining the common frustrations associated with traditional CI/CD pipelines. Relying heavily on YAML or other declarative formats for defining pipelines can lead to complex, repetitive, and hard-to-read configurations, especially for intricate workflows. Debugging failures within these pipelines is often challenging, requiring pushing changes to a remote CI server and waiting for the pipeline to run. Furthermore, pipelines written for one CI platform (like GitHub Actions or GitLab CI) are often not easily transferable to another, creating vendor lock-in and hindering flexibility. This dependency on specific platforms and the difficulty in managing complex workflows manually are significant pain points for development and DevOps teams.

Dagger: CI/CD as Code

Dagger offers a fundamentally different approach by treating CI/CD pipelines as code. It allows developers to write their pipeline logic using familiar programming languages (like Go, Python, Java, or TypeScript) instead of platform-specific configuration languages. This brings the benefits of software development practices – such as code reusability, modularity, testing, and versioning – to CI/CD. Jean-Christophe explained that Dagger executes these pipelines using containers, ensuring consistency and portability across different environments. The Dagger engine runs the pipeline logic, orchestrates the necessary container operations, and manages dependencies. This allows developers to run and debug their CI/CD pipelines locally using the same code that will execute on the remote CI platform, significantly accelerating the debugging cycle.

Dagger Functions and Modules

Key to Dagger’s power are Dagger Functions and Dagger Modules. Jean-Christophe described Dagger Functions as the basic building blocks of a pipeline – functions written in a programming language that perform specific CI/CD tasks (e.g., building a Docker image, running tests, deploying an application). These functions interact with the Dagger engine to perform container operations. Dagger Modules are collections of related Dagger Functions that can be packaged and shared. Modules allow teams to create reusable components for common CI/CD patterns or specific technologies, effectively creating a library of CI/CD capabilities. For example, a team could create a “Java Build Module” containing functions for compiling Java code, running Maven or Gradle tasks, and building JAR or WAR files. These modules can be easily imported and used in different projects, promoting standardization and reducing duplication across an organization’s CI/CD workflows. Jean-Christophe demonstrated how to create and use Dagger Modules, illustrating their potential for building composable and maintainable pipelines. He highlighted that Dagger’s language independence means that modules can be written in one language (e.g., Python) and used in a pipeline defined in another (e.g., Java), fostering collaboration between teams with different language preferences.

The Benefits: Composable, Maintainable, Portable

By adopting Dagger, teams can create CI/CD pipelines that are:
– Composable: Pipelines can be built by combining smaller, reusable Dagger Modules and Functions.
– Maintainable: Pipelines written as code are easier to read, understand, and refactor using standard development tools and practices.
– Portable: Pipelines can run on any platform that supports Dagger and containers, eliminating vendor lock-in.
– Testable: Individual Dagger Functions and modules can be unit tested, and the entire pipeline can be run and debugged locally.

Jean-Christophe’s presentation positioned Dagger as a versatile tool that modernizes CI/CD by bringing the best practices of software development to pipeline automation. The ability to write pipelines in code, leverage reusable modules, and execute locally makes Dagger a powerful “Swiss Army knife” for developers and DevOps engineers seeking more efficient, reliable, and maintainable CI/CD workflows.

Links: