Posts Tagged ‘CloudComputing’
[AWSReInvent2025] Optimizing AWS Costs: Developer-Centric Tools and Methodologies
Lecturer
Kenneth Walsh is a Senior Technical Evangelist at AWS, specializing in cloud financial management (FinOps) and developer productivity. With a background in software engineering and systems architecture, Kenneth focuses on empowering developers to treat “cost as a first-class citizen” in the software development lifecycle. Stacy McOwan is an AWS Developer Advocate who bridges the gap between high-level architectural decisions and day-to-day coding practices. Stacy is a frequent speaker on serverless efficiency and the application of AI to infrastructure management. Together, they provide a pragmatic guide for developers to identify inefficiencies and automate cost optimization using native AWS tools.
Abstract
For the modern cloud developer, the responsibility for system performance and reliability has expanded to include cost efficiency. As cloud environments scale, manual cost management becomes unsustainable, necessitating the adoption of automated, developer-led optimization practices. This article examines the tools and techniques available on AWS to reduce cloud spend without compromising performance. We delve into the use of Amazon Q Developer for AI-powered architectural recommendations and the Kiro CLI for identifying “low-hanging fruit” in resource utilization. The discussion highlights the transition from reactive cost analysis to a “cost-aware” development culture, where optimization is integrated into the CI/CD pipeline. Through the lens of compute, serverless, and observability, this article provides a blueprint for building fiscally responsible applications that maximize the value of every cloud dollar.
The Shift Toward Cost-Aware Development
Historically, cost management was the domain of the finance department or the infrastructure team. However, in a cloud-native world, the code written by a developer directly impacts the AWS bill. A poorly optimized database query or an oversized Lambda function can lead to significant unnecessary expenditure. Kenneth introduces the concept of “cost as a design constraint,” similar to security or latency. When developers are empowered with the right data, they can make informed trade-offs early in the design phase.
Stacy notes that the primary barrier to optimization is often “visibility and friction.” If finding an expensive resource requires navigating dozens of dashboards, it won’t happen. The goal is to bring cost data into the developer’s natural environment—the IDE and the command line. By making optimization a “feature” of the development process, organizations can foster a culture where efficiency is celebrated and waste is proactively eliminated.
AI-Driven Optimization with Amazon Q Developer
One of the most significant innovations in cloud management is the integration of Generative AI into the optimization workflow. Amazon Q Developer serves as a specialized AI assistant that can analyze a developer’s infrastructure and suggest specific, actionable changes. Kenneth demonstrates how Amazon Q can be used to “right-size” instances by analyzing historical CPU and memory usage patterns.
Beyond simple resource sizing, Amazon Q can provide architectural guidance. For example, it might suggest moving a synchronous process to an asynchronous, event-driven model using Amazon SQS to reduce the “idle time” of compute resources. This level of insight allows developers to not just “pay less for what they have” but to “build better systems that cost less by design.”
'''# Example of using AWS SDK to query for cost-optimization recommendations'''
import boto3
client = boto3.client('support')
def get_cost_recommendations():
response = client.describe_trusted_advisor_check_summaries(
checkIds=['eW927uS9S'] # Example ID for Cost Optimization checks
)
for summary in response['summaries']:
print(f"Check: {summary['name']}, Potential Savings: {summary['hasFindings']}")
get_cost_recommendations()
The Kiro CLI: Automating the Identification of Waste
While AI provides high-level guidance, developers often need tactical tools to find specific instances of waste. The Kiro CLI (Cloud Intelligence Reports) is an open-source tool that allows developers to run “cost audits” directly from their terminal. Stacy explains that Kiro can identify “orphaned” resources—such as unattached EBS volumes, old snapshots, or elastic IPs that are not associated with an instance—which are often the biggest contributors to “invisible” cloud spend.
The power of Kiro lies in its ability to be integrated into automation. By running Kiro as part of a weekly “clean-up” script or as a pre-deployment check, teams can ensure that their environments don’t accumulate technical and financial debt over time. Kenneth emphasizes that “low-hanging fruit” optimization—cleaning up what you aren’t using—should be the first step for any organization looking to reduce its cloud bill.
Serverless and Observability: Efficiency in Action
Serverless technologies like AWS Lambda are inherently cost-efficient because they follow a “pay-for-value” model. However, Stacy warns that even serverless can be wasteful if misconfigured. “Lambda Power Tuning” is a methodology where developers test different memory configurations to find the optimal balance between execution speed and cost. Since Lambda charges based on GB-seconds, doubling the memory can sometimes reduce the cost if it cuts the execution time by more than half.
Observability is another area where costs can spiral. Logging everything at “DEBUG” level in production creates massive CloudWatch bills. The lecturers advocate for “intelligent logging,” where detailed logs are only captured during incidents or for a small percentage of transactions. By using Amazon CloudWatch Logs Insights to analyze logging patterns, developers can identify which log groups are generating the most cost and adjust their retention policies accordingly.
Conclusion: Building a Sustainable Cloud Practice
Cost optimization is not a one-time event; it is a continuous practice that requires the right tools, data, and mindset. Kenneth and Stacy conclude that by leveraging AI assistants like Amazon Q and automation tools like the Kiro CLI, developers can take ownership of their cloud spend without it becoming a burden. The ultimate goal is to build applications that are not just technically sound but also economically sustainable. When cost optimization becomes an integral part of the developer workflow, the focus shifts from “cutting costs” to “optimizing value,” enabling the organization to reinvest those savings into further innovation and growth.
Links:
[DevoxxUK2026] Aspiring Speakers Session: Cloud, Code, and Curiosity
Lecturer
Omodolapo Babatunde (often referred to as Dolapo) serves as a Cloud Architect, Communications Manager, and founder of Unora Studio. His work focuses on building products and services for technical professionals, communities, and organizations, with emphasis on digital innovation, inclusive growth, and knowledge sharing.
Abstract
Omodolapo Babatunde explores the critical yet often overlooked phase between initial coding proficiency and mature systems thinking. Emphasizing intentional curiosity, creative experimentation, systems perspectives, and effective communication, the talk offers practical guidance for navigating career ambiguity and fostering sustainable professional development.
Cultivating Growth Through Intentional Curiosity
Omodolapo begins by prompting reflection: when did you last attempt something uncertain? This question frames his journey from uncertainty about industry directions—cloud computing, AI, quantum technologies—toward proactive engagement. Facing fears by weighing potential gains against losses, he advocates intentional curiosity as the catalyst for advancement.
This curiosity manifests through three pillars: creative experimentation, systems thinking, and communication. Experimentation represents the growth sweet spot where developers encounter roadblocks after basic implementation. Rather than stalling, experimentation drives iteration. Omodolapo documented his cloud learning via a blog titled “Cloud, Code, and Curiosity,” which facilitated knowledge sharing, community building, and unexpected opportunities.
Systems thinking emerges once coding fundamentals solidify. It involves evaluating trade-offs, socio-technical implications, and broader impacts on stakeholders. Solutions must consider not merely technical feasibility but organizational and human dimensions.
Communication completes the triad. Translating complex systems requires articulating ideas clearly, listening actively, questioning constructively, and aligning teams. Writing documentation tests true understanding, while presentation skills ensure ideas gain traction.
Omodolapo reinforces these concepts through Receipts, a career record intelligence application enabling professionals to capture, organize, and retrieve achievements efficiently. He encourages audience engagement with the tool.
Conclusion
The presentation culminates in a call to action: remain curious, execute pending experiments, connect with others, and document progress. Whether launching projects, sending messages, or presenting at conferences, proactive steps generate the most valuable career assets. By carrying people along and embracing uncertainty, technologists position themselves at the forefront of innovation.
Links:
[AWSReInvent2025] From Principles to Practice: Scaling AI Responsibly in the Modern Enterprise
Lecturer
Michael Kearns is an Amazon Scholar specializing in Responsible Artificial Intelligence (AI) science, engineering, and policy at Amazon Web Services (AWS). He is a distinguished Professor of Computer and Information Science at the University of Pennsylvania, where his research focuses on machine learning, algorithmic game theory, and the intersection of technology and ethics. Michael is the co-author of The Ethical Algorithm, a seminal work on incorporating social values into software design. His professional background includes extensive experience in quantitative trading and high-level technology consulting.
Kira is a key representative of the Responsible AI team at Indeed, the world’s leading job site. She leads cross-functional initiatives to build tools, systems, and processes that advance inclusive technology. Her work centers on the development of machine learning systems that prioritize fairness, accountability, and transparency to reduce inequalities in the global hiring landscape.
Abstract
The rapid proliferation of generative artificial intelligence (AI) has necessitated a shift from abstract ethical principles to rigorous, operationalized practices. As organizations transition from experimentation to production-scale AI, they face a complex matrix of risks related to privacy, security, fairness, and transparency. This article explores the “AWS Responsible AI Best Practices Framework” and its real-world application at Indeed. By examining how Indeed has built an intelligent risk management platform, the analysis highlights the necessity of embedding responsibility at every stage of the AI lifecycle. The discussion moves beyond compliance, illustrating how a robust “Responsible AI (RAI) posture” can accelerate innovation by building trust and ensuring enterprise-grade safety.
Introduction to the Responsible AI Lifecycle
The contemporary AI landscape is defined by a tension between the desire for rapid innovation and the imperative to mitigate systemic risks. While the “what” of responsible AI—fairness, safety, and privacy—is well-established, the “how” remains a significant challenge for many enterprises. At AWS, the philosophy of Responsible AI is integrated into the core service architecture, emphasizing that responsibility is not a final checkbox but a continuous process.
Michael identifies that every AI system possesses an inherent “REI posture,” whether intentionally designed or not. This posture is influenced by data selection, model tuning, and deployment context. The AWS framework encourages organizations to move toward “platformization,” where responsible checks are built directly into the developer workflow. This approach ensures that developers do not have to choose between speed and safety; instead, the platform provides the necessary guardrails.
The Indeed Case Study: Embedding Fairness in Hiring
Hiring is a fundamentally human process where the stakes are exceptionally high. For Indeed, the mission is to help people get jobs, making fairness and the reduction of bias central to their technological identity. Kira explains that talent is universal, but opportunity is not. AI has the potential to either dismantle or amplify existing barriers in the job market.
Indeed’s methodology for scaling AI responsibly involves several critical pillars:
- Job Seeker First: All AI development is guided by the ultimate impact on the end-user.
- Multidisciplinary Governance: Indeed utilizes a cross-functional team that bridges the gap between legal requirements, social science, and engineering.
- The Responsible AI Lens: By utilizing tools like the AWS Well-Architected Tool, Indeed evaluates its systems across multiple dimensions of responsibility, including robustness and explainability.
Methodologies for Risk Mitigation and Platformization
The transition from “principles to practice” requires tangible tools. One of the primary innovations discussed is the creation of an intelligent risk management platform. This platform serves as a centralized hub for monitoring how AI products interact with job seekers and employers in real-time.
Anticipatory Guardrails
Before a model reaches production, it must undergo rigorous testing for fairness. Indeed incorporates the “lived experiences” of job seekers into their testing phase, recognizing that quantitative data alone may not capture the nuances of cultural context or systemic bias. By setting up proactive guardrails, the organization can block the deployment of models that do not meet predefined safety and fairness thresholds.
Continuous Monitoring and Feedback
Once a system is live, the work continues. Indeed’s infrastructure is designed for “REI observability.” This involves tracking signals such as log metrics and user traces to detect drift or unintended consequences. Because the definition of “fairness” is highly contextual and evolves over time, Indeed maintains a “listen and learn” journey, iterating on their models based on both data-driven insights and qualitative feedback from the community.
Consequences for Enterprise Strategy
The implications of adopting a comprehensive RAI framework are twofold. First, it satisfies the increasing pressure from global regulators and policymakers. By aligning with frameworks such as the NIST AI Risk Management Framework, companies like Indeed and AWS stay ahead of legislative mandates.
Second, and perhaps more importantly, responsible AI acts as a business differentiator. In an era where consumer trust is fragile, demonstrating a commitment to transparency and safety builds long-term brand loyalty. Michael emphasizes that by building a “box” or a “sandbox” for agents and models that is secure and observable, organizations actually unlock their development teams. When developers know they are playing in a safe environment, they are more willing to experiment with production-grade tools and real customer data.
Conclusion
Scaling AI responsibly is no longer an optional ethical exercise; it is a foundational requirement for production-grade engineering. The journey from high-level principles to operational practice involves the integration of cross-functional expertise, the deployment of specialized risk-management platforms, and a culture of continuous learning. As demonstrated by the collaboration between AWS and Indeed, the future of AI belongs to those who can build systems that are not only powerful but also trusted, transparent, and fair.
Links:
[AWSReInvent2025] Maximizing Block Storage Performance for High-Intensity Workloads: A Technical Analysis of io2 Block Express and the Nitro System
Lecturer
Mark Olsen and Jody Berenblatt are distinguished engineering and product leaders at Amazon Web Services, specializing in high-performance block storage. Mark Olsen serves as a Principal Product Manager for Amazon EBS, where he focuses on the architectural evolution of Provisioned IOPS volumes to meet the demands of mission-critical enterprise applications. Jody Berenblatt, a Senior Technical Product Manager, brings extensive expertise in the integration of storage subsystems with the AWS Nitro System and the optimization of storage networking protocols. Their work has been pivotal in the development of io2 Block Express, a storage tier designed to provide SAN-like performance in the cloud.
Abstract
This article provides a comprehensive examination of the technical foundations and performance characteristics of high-intensity block storage within the Amazon Elastic Block Store (EBS) ecosystem. Centered on the io2 Block Express architecture, the analysis explores how the integration of the AWS Nitro System, the Scalable Reliable Datagram (SRD) protocol, and Multi-Attach NVMe reservations enables ultra-low latency and high-throughput capabilities for data-intensive workloads such as SAP HANA, Oracle, and Microsoft SQL Server. The discussion details the methodology for managing tail latency, the benefits of decoupled storage architectures, and the operational strategies required to maximize I/O performance in a distributed cloud environment.
Infrastructure Foundations: The Evolution of Provisioned IOPS
The landscape of enterprise computing has shifted toward workloads that demand not only high throughput but also extreme consistency in I/O operations per second (IOPS). For decades, on-premises Storage Area Networks (SANs) were the only viable option for these applications. However, the maturation of Amazon EBS, particularly the transition from io1 to the io2 Block Express architecture, has redefined the capabilities of cloud-native block storage. The fundamental challenge in high-intensity storage is the management of latency, which is often the primary bottleneck for database performance.
In traditional storage models, performance was often tethered to the physical limitations of the disk or the controller. In the modern AWS architecture, the storage is decoupled from the compute instance, connected via a dedicated high-speed network. This separation allows for independent scaling of compute and storage resources but introduces the necessity for highly optimized networking to maintain sub-millisecond latency. The io2 Block Express volumes are engineered to provide up to 256,000 IOPS and 4,000 MB/s of throughput per volume, offering a level of performance that satisfies even the most demanding transactional databases.
Architecture of io2 Block Express: Performance and Durability
The architecture of io2 Block Express represents a paradigm shift in how block storage is provisioned and managed. Unlike standard volumes, io2 Block Express is designed to handle “high-intensity” workloads, defined by their sensitivity to latency and their requirement for high durability. These volumes provide a durability rating of 99.999%, which is a ten-fold improvement over standard io1 volumes. This reliability is achieved through sophisticated replication techniques across multiple physical hardwares within an Availability Zone.
A critical innovation in this architecture is the way it handles I/O operations. By utilizing the Nitro System, the overhead of the hypervisor is removed, allowing the EBS service to communicate directly with the instance’s memory. This “Block Express” layer acts as a high-performance interface that minimizes the processing time required for each I/O request. For applications like SAP HANA, where the speed of logging and data loading is critical, the reduced overhead translates directly into faster business processing cycles.
Networking Innovations: Scalable Reliable Datagram (SRD)
Perhaps the most significant technical advancement in maximizing block storage performance is the implementation of the Scalable Reliable Datagram (SRD) protocol. Traditional TCP protocols, while reliable, are prone to “head-of-line blocking,” where a single lost packet can delay the entire stream of data. In a high-performance storage environment, this creates “tail latency”—spikes in response time that can disrupt database synchronization and performance.
SRD solves this by utilizing multipath routing. Instead of sending data down a single network path, SRD spreads the traffic across as many as 64 different paths simultaneously. If a specific network switch becomes congested or a link fails, the protocol automatically reroutes the data without the latency spikes associated with TCP retransmissions. This protocol is implemented directly in the Nitro Cards, ensuring that the heavy lifting of network management does not consume CPU cycles on the user’s EC2 instance. The result is a more consistent “p99” latency profile, which is essential for maintaining stable performance in clustered environments.
Multi-Attach NVMe Reservations and High Availability
For enterprise applications requiring high availability, the ability for multiple EC2 instances to attach to a single EBS volume is a critical requirement. io2 Block Express supports Multi-Attach, allowing up to 16 Nitro-based instances to access the same volume simultaneously. This feature is particularly valuable for clustered file systems and applications that require shared storage for failover or parallel processing.
To manage concurrent access without data corruption, AWS implemented Multi-Attach NVMe Reservations. Based on the NVMe standard for persistent reservations (similar to SCSI-3 PR), this technology allows one instance to “reserve” the volume, ensuring that only authorized nodes can perform write operations. In the event of an instance failure, the reservation can be quickly cleared and reassigned to a healthy node, minimizing downtime. This mechanism provides the coordination layer necessary for complex deployments like Oracle RAC or SAP environments, where data integrity across multiple nodes is non-negotiable.
Observability and Performance Tuning for Enterprise Workloads
Achieving maximum performance requires a sophisticated approach to observability. Many administrators focus on average latency, but in high-intensity workloads, the “outliers” or tail latency are what truly matter. AWS provides tools such as Amazon CloudWatch and EBS Volume Insights to monitor these metrics in real-time. A key metric is the “Queue Depth,” which represents the number of pending I/O requests for a volume. To reach the full potential of an io2 Block Express volume (e.g., 256,000 IOPS), the application must maintain a sufficient queue depth—often 128 or higher—to keep the storage pipeline full.
// Example AWS CLI command to modify an EBS volume to io2 with high provisioned IOPS
aws ebs modify-volume \
--volume-id vol-0123456789abcdef \
--volume-type io2 \
--iops 100000
Furthermore, the choice of the EC2 instance type is paramount. Performance is not solely a function of the storage volume; the instance must be “EBS-optimized” with sufficient dedicated bandwidth to handle the provisioned throughput. For instance, using an R5b or X2idn instance allows the application to utilize the full 4,000 MB/s throughput offered by Block Express. Failure to match the instance capability with the volume performance will lead to throttling at the instance level, regardless of how many IOPS are provisioned.
Links:
[AWSReInvent2025] Beyond Migration: Transforming Global Automotive Retail with SAP and Pan-Amazon Services
Lecturer
Sunnuk Kim is the Vice President and Head of the IT Strategy and Planning Division at Hyundai Motor Group. Based in Seoul, he is a primary architect of the group’s digital strategy, focusing on integrating legacy industrial operations with modern cloud intelligence to redefine the automotive lifecycle. Mahesh Shrivastava is a Director and Global Leader for SAP on AWS. He specializes in enterprise-scale digital transformation, helping multinational corporations move beyond infrastructure optimization to achieve true business model innovation through cloud-native ecosystems.
Abstract
The modern enterprise technology landscape is undergoing a fundamental shift where global organizations no longer view cloud migration as an isolated technical objective but rather as a catalyst for comprehensive business transformation. This article examines the strategic collaboration between Hyundai Motor Group and Amazon Web Services (AWS) to modernize its mission-critical SAP environment through the integration of “Pan-Amazon” services. By moving beyond traditional “lift-and-shift” methodologies, Hyundai has adopted a “clean core” strategy that bridges the gap between back-office ERP functions and front-end consumer touchpoints. The analysis explores how the integration of Amazon Business, Prime logistics, and multi-channel fulfillment centers with SAP allows Hyundai to optimize global sales, inventory management, and personalized retail experiences. This transformation signifies the evolution of the automotive industry into a data-driven, customer-centric retail model.
The Strategic Shift: From Infrastructure Migration to Business Evolution
Historically, large-scale enterprises approached the cloud with the narrow objective of reducing capital expenditure by transitioning physical data centers to virtualized environments. For a global manufacturer like Hyundai, the initial focus was often on the stability and performance of SAP systems that manage the “heartbeat” of production and finance. However, as market dynamics evolved toward direct-to-consumer models and digital-first interactions, the group identified that true value lay in how cloud-native capabilities could solve complex business challenges. This realization prompted a move away from simply “running” SAP in the cloud toward “transforming” the business through the cloud.
The strategic pivot was driven by an urgent need for customer-centricity, requiring Hyundai to provide seamless, omnichannel experiences that mirror the speed and predictability of modern e-commerce. Furthermore, the limitations of rigid, monolithic legacy architectures necessitated a “clean core” approach. This methodology allows the organization to maintain a stable, standard ERP foundation while rapidly innovating through extensions and external integrations. By breaking down the long-standing silos between manufacturing data and external consumer insights, Hyundai has positioned itself to make real-time decisions that directly impact global sales volume and customer retention.
Methodology: Integration of the Pan-Amazon Ecosystem
A core innovation in Hyundai’s transformation is the sophisticated utilization of “Pan-Amazon” services, a broad collection of Amazon’s diverse business units that are now integrated directly into the AWS cloud platform. This strategy extends far beyond typical compute and storage services. For instance, the integration of Amazon Business has allowed Hyundai to streamline indirect procurement and supply chain management directly within the SAP workflow, reducing manual overhead and improving spend visibility.
Furthermore, the application of Amazon Prime and its global fulfillment network to the automotive sector represents a significant methodology shift. By leveraging these world-class logistics models, Hyundai can manage automotive parts and vehicle accessories with unprecedented efficiency. This creates a “Y process” where product portfolio management and sales volume planning converge. In this model, the back-office operations managed by SAP are directly connected to the front-end retail experience. This integration ensures that when a customer interacts with a digital retail channel, the system can provide real-time data on vehicle availability, delivery timelines, and personalized configuration options, all backed by a robust, cloud-native logistics engine.
Technical Analysis of Modernized Operations
The transition from legacy environments to an AWS-integrated SAP landscape has yielded transformative results across several key performance indicators. In terms of scalability, the previous architecture was constrained by fixed capacity and physical hardware limitations, whereas the current AWS-integrated system offers elastic scaling that adapts to real-time demand spikes without manual intervention. Global inventory management has transitioned from fragmented data silos, which often suffered from latency and inaccuracies, to a unified system providing real-time visibility across all global fulfillment centers.
Customer experience has seen a similar leap in sophistication. What was once a linear and offline-heavy journey has been replaced by an integrated omnichannel digital retail platform that provides consumers with the speed and reliability they expect from modern digital platforms. This operational efficiency at scale is further supported by the ability to access the world’s largest online marketplace and fulfillment network. The technical result is a modular environment where the core ERP remains upgradable and stable while a vast array of custom, cloud-native services drive innovation on the periphery. This architecture ensures that even as the company expands into new geographic regions or business channels, the underlying infrastructure remains resilient and performant.
Implications for Global Automotive Retail and Beyond
The consequences of Hyundai’s “Go to Cloud” strategy are profound for the broader automotive sector. The industry is moving toward a state of direct-to-consumer readiness, where traditional dealership models are being augmented by digital platforms that offer complete transparency and predictability. This shift is enabled by the ability to treat vehicle sales not as a one-time transaction, but as a continuous relationship supported by digital services and efficient parts logistics.
The success of this project also highlights the importance of data-driven innovation. By analyzing vast amounts of data across the combined SAP and Amazon ecosystem, Hyundai can better forecast market trends and optimize production cycles accordingly. This represents a broader trend of Industry 4.0, where the lines between manufacturing, retail, and technology are increasingly blurred. The ability to achieve such high levels of operational agility while maintaining a secure and compliant global footprint sets a new benchmark for enterprise-scale digital transformation.
Conclusion
The collaboration between Hyundai Motor Group and AWS serves as a comprehensive blueprint for how large enterprises can successfully navigate the complexities of modernizing mission-critical systems. By prioritizing the customer experience and leveraging the full breadth of the Pan-Amazon ecosystem, Hyundai has evolved from a traditional manufacturer into a leader in digital automotive retail. The journey underscores that the future of enterprise IT is defined not just by the technology itself, but by the intelligent integration of diverse services to create tangible business value. As global competition intensifies, the move toward a “clean core” SAP environment supported by cloud-native logistics and AI will be the defining factor for sustainable growth and innovation.
Links:
[MunchenJUG] Navigating the JVM Ecosystem: A Safari Through Distributions (16/Sep/2024)
Lecturer
Gerrit Grunwald is a highly regarded software engineer and advocate with four decades of experience in the technology sector. He is a prominent figure in the Java community, recognized as a Java Champion and a JavaOne Rockstar. Gerrit is deeply committed to open-source software, having contributed to and led numerous projects such as JFXtras, TilesFX, Medusa, and JDKMon. He founded and leads the Java User Group Münster and is a frequent speaker at international conferences. Currently, Gerrit serves as a Developer Advocate at Azul.
Abstract
This article provides an analytical overview of the modern Java Virtual Machine (JVM) landscape, distinguishing between the OpenJDK project and its various commercial and community distributions. It evaluates the shift in Java’s release cadence and the implications for long-term support (LTS) in corporate environments. A significant portion of the analysis is dedicated to the optimization of Java runtimes through modularity and the jlink tool, demonstrating how developers can significantly reduce deployment sizes and enhance security. Finally, the article categorizes the plethora of available JDK distributions—from major cloud providers like Amazon and Alibaba to specialized runtimes like GraalVM—offering a guide for selecting the appropriate distribution based on specific use cases.
The Distinction Between OpenJDK and Distributions
A fundamental misunderstanding in the Java community is the conflation of “OpenJDK” with the software installed on a user’s machine. OpenJDK is not a downloadable product but rather the open-source project hosted on GitHub that contains the source code for the Java Platform, Standard Edition (Java SE). What developers actually utilize are “builds” or “distributions” of this source code.
The OpenJDK ecosystem is characterized by its collaborative nature, with significant contributions from tech giants such as Oracle, Amazon, ARM, Google, Intel, and IBM. This multi-corporate backing ensures the longevity and stability of the platform, preventing it from becoming a “one-man show”. Since moving to GitHub with JDK 16, the transparency and accessibility of the source code have further improved, allowing for faster build times and broader community involvement.
Release Cadence and Support Models
The evolution of Java’s release model marks a critical transition from multi-year development cycles to a predictable six-month cadence. Historically, long gaps between releases (such as the five years between JDK 6 and JDK 7) led to massive, overwhelming updates that were difficult for organizations to adopt.
The current model classifies releases into two categories:
- Feature Releases: Released every six months, these versions typically receive support for only half a year.
- Long-Term Support (LTS) Releases: These versions are designated for extended support, often spanning a decade or more, providing the stability required by enterprise applications.
This dual-track approach allows the language to innovate rapidly through feature releases while providing a safe harbor for production environments on LTS versions.
Efficiency through Modularity: The jlink Revolution
One of the most underutilized innovations introduced in JDK 9 is the modularization of the Java runtime. By breaking the monolithic JDK into 69 distinct modules, Oracle enabled developers to create custom, stripped-down runtimes tailored to specific applications.
The tool jlink allows for the creation of a custom Java Runtime Environment (JRE) containing only the modules necessary for a particular application. The impact on deployment size is profound:
- A full JDK 21 installation requires approximately 340 MB.
- A standard JRE for the same version takes about 150 MB.
- A
jlink-optimized runtime for a simple application (like a push notification server) can be as small as 48 MB.
echo Example of using jdeps to find required modules
jdeps --ignore-missing-deps --print-module-deps MyProject.jar
echo Example of using jlink to create a custom runtime
jlink --add-modules java.base,java.logging --output custom-runtime
Beyond storage savings, modular runtimes enhance security by reducing the attack surface. If a vulnerability exists in a module that has been excluded from the custom runtime (such as the desktop module in a server-side application), the application remains unaffected.
Mapping the Distribution Jungle
The JVM landscape is populated by numerous distributions, each offering different levels of support, licensing, and platform optimizations.
Community and Vendor Builds
- Eclipse Temurin (formerly AdoptOpenJDK): A widely used community build that is TCK (Technology Compatibility Kit) compliant.
- Amazon Corretto: A no-cost, multiplatform distribution used internally by Amazon for its AWS services.
- Azul Zulu: A TCK-compliant distribution offering broad platform support.
- Oracle OpenJDK: The free, GPL-licensed build provided by Oracle.
Region-Specific and Specialized Distributions
In the Asian market, distributions like Alibaba’s Dragonwell, Huawei’s Bi Sheng, and Tencent’s Kona are dominant. These often include specific optimizations for the cloud infrastructures of their respective parent companies.
Advanced Runtimes: GraalVM and Beyond
GraalVM represents a specialized branch of the JVM ecosystem, offering high-performance polyglot capabilities and “Native Image” compilation. Native images allow Java applications to start in milliseconds by compiling them into platform-specific executables, though this comes at the cost of peak performance and longer build times compared to the standard JIT (Just-In-Time) compilation used by the HotSpot JVM.
Conclusion: Strategy for Selection
Choosing the right JVM distribution is a strategic decision based on support requirements, cost, and technical constraints. For most production environments, sticking to an LTS version from a reputable vendor (like Azul, Amazon, or the Eclipse Foundation) ensures stability. Meanwhile, developers should leverage modern tools like jlink to ensure their deployments remain lean and secure, regardless of the distribution chosen.
Links:
[AWSReInvent2025] Supercharging DevOps with AI-Driven Observability: The Next Frontier in SRE
Lecturer
Elizabeth Fuentes is a Senior Developer Advocate at Amazon Web Services (AWS), specializing in the intersection of Artificial Intelligence and DevOps practices. With extensive experience in cloud architecture and software engineering, Elizabeth focuses on how Generative AI can streamline complex CI/CD pipelines and enhance Site Reliability Engineering (SRE). She is a key contributor to AWS educational initiatives, having co-developed advanced courses on AI-driven automation. Joining her is Laas Alina, a software architect and open-source enthusiast who focuses on implementing multi-agent systems and the Model Context Protocol (MCP) to solve observability challenges at scale.
Abstract
As software systems grow increasingly distributed and complex, traditional observability—centered on manual log analysis and reactive dashboards—is becoming insufficient. This article explores the paradigm shift toward AI-driven observability, where Generative AI serves not just as a query tool, but as an active participant in failure detection, correlation, and resolution. By leveraging Amazon Bedrock and Amazon Q, organizations can transition from “reactive” to “predictive” DevOps. The discussion analyzes the methodology of building AI agents that simulate architectural stress, automatically explain multi-layered failures, and provide traceable, actionable recommendations. We examine the implementation of the Model Context Protocol (MCP) in establishing sophisticated multi-agent systems (MAS) that transform raw data into contextual understanding, ultimately reducing the Mean Time to Resolution (MTTR) and enhancing systemic resilience.
The Evolution of Observability: From Metrics to Contextual Understanding
The traditional pillars of observability—metrics, logs, and traces—provide the “what” of a system’s state but often fail to provide the “why” in real-time. In high-velocity DevOps environments, the sheer volume of telemetry data can overwhelm human operators, leading to “alert fatigue” and delayed responses to critical incidents. Elizabeth posits that the integration of Generative AI marks the fourth pillar of observability: Contextual Intelligence. This evolution moves the industry beyond simple threshold-based monitoring toward systems that understand the semantic relationship between a failed deployment, a spike in latency, and a specific line of code.
By utilizing Large Language Models (LLMs) through Amazon Bedrock, DevOps teams can ingest vast amounts of unstructured log data and receive summaries that highlight anomalies that might be missed by traditional regex-based filters. The methodology involves training the AI to recognize “normal” operational patterns and identifying deviations not just by value, but by the intent of the system’s behavior. This contextual layer allows for a more nuanced interpretation of system health, where the AI can distinguish between a benign resource spike and a precursor to a cascading failure.
Architecting AI Agents for Predictive Troubleshooting
The transition to AI-driven observability is characterized by the deployment of “Micro-agents”—specialized AI entities designed to handle specific segments of the DevOps lifecycle. These agents operate within a Multi-Agent System (MAS), where they collaborate to solve complex incidents. For instance, a “Monitoring Agent” might detect a performance degradation and immediately trigger a “Diagnosis Agent” to correlate the event with recent CI/CD pipeline changes.
Elizabeth and Laas Alina emphasize the importance of the Model Context Protocol (MCP) in this architecture. MCP acts as the communication backbone, allowing agents to share context without losing the “lineage” of a decision. When an AI agent recommends a specific architectural change or a rollback, it must provide clear traceability. This is crucial for maintaining trust in automated systems. The agents do not operate in a vacuum; they interact with tools like Amazon Q to provide developers with instant explanations of failures directly within their Integrated Development Environment (IDE) or chat interface.
// Example of an AI-driven Observability Agent Configuration
agent:
name: "IncidentDiagnosticAgent"
provider: "AmazonBedrock"
model: "claude-3-sonnet"
capabilities:
- log_analysis
- metric_correlation
- trace_summarization
mcp_config:
protocol_version: "1.0"
shared_context: "deployment_metadata"
safety_guardrails:
- max_token_usage: 4000
- human_in_the_loop_required: true
Transforming CI/CD through Generative AI and Simulation
Beyond reactive troubleshooting, AI-driven observability empowers proactive system design. One of the most innovative concepts discussed is the use of AI agents to simulate “stress-test” scenarios within a digital twin of the production environment. These agents can intentionally inject failures—similar to Chaos Engineering—and then observe how the observability stack responds. This creates a feedback loop where the AI helps engineers identify “blind spots” in their monitoring before a real incident occurs.
Furthermore, Generative AI transforms the CI/CD pipeline by automatically generating “failure explanations.” Instead of a developer sifting through a 5,000-line build log, Amazon Q can provide a concise summary: “The build failed because the new database schema in commit X is incompatible with the connection pool settings in environment Y.” This level of automated insight accelerates the “inner loop” of development, allowing engineers to focus on innovation rather than infrastructure archeology.
The Human-AI Partnership: Strategic Implications
A common concern in the industry is the replacement of human engineers by AI. However, Elizabeth argues that the future belongs to the “augmented engineer.” AI is a force multiplier that automates the repetitive, “drudge work” of observability—log parsing and initial triage—allowing human experts to focus on high-level strategy and complex architectural decisions. The goal is to transform teams from being “reactive” (fighting fires) to “proactive” (preventing fires).
Implementing these systems requires a cultural shift toward AI-literacy within DevOps teams. Organizations must establish safety guardrails to ensure that AI-driven recommendations are validated and that automated actions (like auto-remediation) have clear rollback paths. By embracing AI as a strategic tool, DevOps and SRE teams can achieve a level of operational excellence that was previously unattainable, ensuring that as systems grow in scale, their reliability grows in parallel.
Links:
[AWSReInvent2025] High-Performance Storage Architectures for AI/ML, Analytics, and HPC Workloads
Lecturer
Aditi is a Senior Product Manager for Amazon FSx at Amazon Web Services (AWS). With years of experience working directly with customers on high-performance workloads, she focuses on pushing the technical boundaries of what is possible with cloud storage to meet the demands of modern compute-intensive applications.
Abstract
This article examines the critical role of high-performance storage in supporting modern AI/ML, analytics, and High-Performance Computing (HPC) workloads. As organizations scale their compute resources—incorporating hundreds or thousands of CPU and GPU cores—storage often becomes the primary bottleneck, preventing linear performance scaling. We explore the technical architectures of Amazon FSx and Amazon S3, focusing on how these services address the needs of both “lift-and-shift” file-based applications and “cloud-native” S3-based data lakes. By analyzing customer use cases in genomics, media rendering, and large language model (LLM) training, we detail the methodologies for achieving peak performance at scale.
The Storage Bottleneck in Compute-Intensive Workloads
Modern high-performance workloads are characterized by their extreme reliance on massive datasets and high-core-count compute clusters. In an ideal cloud environment, adding more compute resources should lead to a proportional increase in work completed—a concept known as linear scaling. However, traditional storage solutions often fail to keep pace with the throughput demands of these clusters, leading to a performance plateau.
When storage becomes the bottleneck, compute instances sit underutilized as they compete for access to the same data store. This is particularly detrimental given that 90% to 95% of the expenditure for these workloads is typically allocated to compute resources. Consequently, an inefficient storage layer not only extends the time to insight but also significantly increases the total cost of ownership (TCO). To avoid this, storage must be architected to scale linearly alongside compute.
Navigating the Path to the Cloud: File Systems vs. Object Storage
Organizations generally approach high-performance storage on AWS from two distinct backgrounds: those with long-standing on-premises file-based workflows and those who have built native cloud applications around object storage.
The Persistence of File-Based Architectures
Despite the rise of object storage, file systems remain the preferred interface for many researchers and developers due to three primary factors: Familiar Interface: The intuitive nature of files and directories simplifies complex data management for data scientists and developers.
* Granular Permissions: File systems provide robust POSIX permissions, allowing for fine-grained control over which users can read, write, or execute specific files.
* Consistent Data Access:* For workloads where multiple users or compute nodes access the same data simultaneously, the strong consistency of file systems ensures that all parties see the most recent data updates.
Amazon FSx for High-Performance File Access
Amazon FSx addresses these needs by providing fully managed file systems that offer the performance of local storage with the scalability of the cloud. For “lift-and-shift” scenarios, FSx allows organizations to move their existing HPC and AI/ML pipelines to AWS without refactoring their applications.
Accelerating Generative AI and ML Workloads
The emergence of generative AI has placed a renewed emphasis on data strategy. Whether an organization is building a model from scratch or fine-tuning a foundational model, the quality and accessibility of its proprietary data are the primary differentiators.
Retrieval Augmented Generation (RAG)
To move beyond generic AI responses and reduce hallucinations, many organizations are implementing Retrieval Augmented Generation (RAG). RAG allows foundational models to access evolving, large-scale data lakes without requiring the data to be manually loaded into a prompt.
The RAG methodology involves:
1. Vectorization: Converting organizational data into vectors—numeric representations that capture semantic meaning.
2. Semantic Search: Using spatial similarity to compare a query vector against the data lake’s vectors to find the most relevant information.
3. Augmentation: Feeding the retrieved context back into the model to generate a more accurate and business-specific response.
Ingestion and Data Strategy with Amazon S3
Amazon S3 serves as the foundational data lake for these AI workflows due to its cost-effectiveness and virtually unlimited scalability. Organizations typically utilize two ingestion patterns:
* Batch Ingestion: Suitable for static or infrequently changing data such as historical records and product catalogs.
* Real-Time Ingestion: Essential for agentic workflows where AI models must respond to the latest available information.
Modernizing Self-Managed Databases with Amazon FSx
While fully managed services like Amazon RDS are popular, certain business and technical requirements drive organizations toward self-managed database architectures on AWS.
Drivers for Self-Managed Databases
Organizations choose to self-manage databases like Oracle, SQL Server, or SAP HANA for several reasons:
* Granular Control: The ability to choose specific versions of the database engine and the underlying operating system.
* Custom Protection Policies: Implementing specific backup intervals and recovery procedures that may not be available in managed services.
* High Resilience: Scaling databases across multiple Availability Zones or regions with custom failover configurations.
Optimization through Storage Features
A common oversight in database deployment is the potential for the storage layer to add significant value beyond simple data persistence. Amazon FSx file systems (including FSx for NetApp ONTAP, OpenZFS, and Windows File Server) enable features like:
* Snapshots and Cloning: Facilitating rapid testing and database upgrades by creating near-instantaneous copies of production environments.
* Performance Tuning: Choosing the right FSx service can significantly optimize the TCO and performance of database environments, particularly for high-transaction workloads.
Conclusion
As compute power continues to expand, the storage layer must evolve from a passive repository into a high-performance engine. By leveraging Amazon FSx and S3, organizations can eliminate storage bottlenecks, enabling their most demanding AI, HPC, and database workloads to scale linearly and cost-effectively in the cloud.
Links:
- Amazon FSx Product Page
- Amazon S3 Product Page
- AWS re:Invent 2025 – High-performance storage for AI/ML, analytics, and HPC workloads (STG336)
- AWS re:Invent 2025 – Accelerate gen AI and ML workloads with AWS storage (STG201)
- AWS re:Invent 2025 – Improve self-managed database performance and agility with Amazon FSx (STG337)
[DevoxxPL2022] Challenges Running Planet-Wide Computer: Efficiency • Jacek Bzdak, Beata Strack
Jacek Bzdak and Beata Strack, software engineers at Google Poland, delivered an engaging session at Devoxx Poland 2022, exploring the intricacies of optimizing Google’s planet-scale computing infrastructure. Their talk focused on achieving efficiency in a distributed system spanning global data centers, emphasizing resource utilization, auto-scaling, and operational strategies. By sharing insights from Google’s internal cloud and Autopilot system, Jacek and Beata provided a blueprint for enhancing service performance while navigating the complexities of large-scale computing.
Defining Efficiency in a Global Fleet
Beata opened by framing Google’s data centers as a singular “planet-wide computer,” where efficiency translates to minimizing operational costs—servers, CPU, memory, data centers, and electricity. Key metrics like fleet-wide utilization, CPU/RAM allocation, and growth rate serve as proxies for these costs, though they are imperfect, often masking quality issues like inflated memory usage. Beata stressed that efficiency begins at the service level, where individual jobs must optimize resource consumption, and extends to the fleet through an ecosystem that maximizes resource sharing. This dual approach ensures that savings at the micro level scale globally, a principle applicable even to smaller organizations.
Auto-Scaling: Balancing Utilization and Reliability
Jacek, a member of Google’s Autopilot team, delved into auto-scaling, a critical mechanism for achieving high utilization without compromising reliability. Autopilot’s vertical scaling adjusts resource limits (CPU/memory) for fixed replicas, while horizontal scaling modifies replica counts. Jacek presented data from an Autopilot paper, showing that auto-scaled services maintain memory slack below 20% for median cases, compared to over 60% for manually managed services. Crucially, automation reduces outage risks by dynamically adjusting limits, as demonstrated in a real-world case where Autopilot preempted a memory-induced crash. However, auto-scaling introduces complexity, particularly feedback loops, where overzealous caching or load shedding can destabilize resource allocation, requiring careful integration with application-specific metrics.
Java-Specific Challenges in Auto-Scaling
The talk transitioned to language-specific hurdles, with Jacek highlighting Java’s unique challenges in auto-scaling environments. Just-in-Time (JIT) compilation during application startup spikes CPU usage, complicating horizontal scaling decisions. Memory management poses further issues, as Java’s heap size is static, and out-of-memory errors may be masked by garbage collection (GC) thrashing, where excessive CPU is devoted to GC rather than request handling. To address this, Google sets static heap sizes and auto-scales non-heap memory, though Jacek envisioned a future where Java aligns with other languages, eliminating heap-specific configurations. These insights underscore the need for language-aware auto-scaling strategies in heterogeneous environments.
Operational Strategies for Resource Reclamation
Beata concluded by discussing operational techniques like overcommit and workload colocation to reclaim unused resources. Overcommit leverages the low probability of simultaneous resource spikes across unrelated services, allowing Google to pack more workloads onto machines. Colocating high-priority serving jobs with lower-priority batch workloads enables resource reclamation, with batch tasks evicted when serving jobs demand capacity. A 2015 experiment demonstrated significant machine savings through colocation, a concept influencing Kubernetes’ design. These strategies, combined with auto-scaling, create a robust framework for efficiency, though they demand rigorous isolation to prevent interference between workloads.
Links:
[PHPForumParis2021] Migrating a Bank-as-a-Service to Serverless – Louis Pinsard
Louis Pinsard, an engineering manager at Theodo, captivated the Forum PHP 2021 audience with a detailed recounting of his journey migrating a Bank-as-a-Service platform to a serverless architecture. Having returned to PHP after a hiatus, Louis shared his experience leveraging AWS serverless technologies to enhance scalability and reliability in a high-stakes financial environment. His narrative, rich with practical insights, illuminated the challenges and triumphs of modernizing critical systems. This post explores four key themes: the rationale for serverless, leveraging AWS tools, simplifying with Bref, and addressing migration challenges.
The Rationale for Serverless
Louis Pinsard opened by explaining the motivation behind adopting a serverless architecture for a Bank-as-a-Service platform at Theodo. Traditional server-based systems struggled with scalability and maintenance under the unpredictable demands of financial transactions. Serverless, with its pay-per-use model and automatic scaling, offered a solution to handle variable workloads efficiently. Louis highlighted how this approach reduced infrastructure management overhead, allowing his team to focus on business logic and deliver a robust, cost-effective platform.
Leveraging AWS Tools
A significant portion of Louis’s talk focused on the use of AWS services like Lambda and SQS to build a resilient system. He described how Lambda functions enabled event-driven processing, while SQS managed asynchronous message queues to handle transaction retries seamlessly. By integrating these tools, Louis’s team at Theodo ensured high availability and fault tolerance, critical for financial applications. His practical examples demonstrated how AWS’s native services simplified complex workflows, enhancing the platform’s performance and reliability.
Simplifying with Bref
Louis discussed the role of Bref, a PHP framework for serverless applications, in streamlining the migration process. While initially hesitant due to concerns about complexity, he found Bref to be a lightweight layer over AWS, making it nearly transparent for developers familiar with serverless concepts. Louis emphasized that Bref’s simplicity allowed his team to deploy PHP code efficiently, reducing the learning curve and enabling rapid development without sacrificing robustness, even in a demanding financial context.
Addressing Migration Challenges
Concluding his presentation, Louis addressed the challenges of migrating a legacy system to serverless, including team upskilling and managing dependencies. He shared how his team adopted AWS CloudFormation for infrastructure-as-code, simplifying deployments. Responding to an audience question, Louis noted that Bref’s minimal overhead made it a viable choice over native AWS SDKs for PHP developers. His insights underscored the importance of strategic planning and incremental adoption to ensure a smooth transition, offering valuable lessons for similar projects.