Recent Posts
Archives

Posts Tagged ‘GenerativeAI’

PostHeaderIcon [DevoxxFR2026] Measuring the Unmeasurable: Evaluating Generative AI Systems

Lecturer

Erin Pacquetet is an expert in AI evaluation and product development at SCIAM, a Paris-based consulting firm. With a background in linguistics and extensive experience guiding enterprises through the complexities of deploying generative AI applications, she specializes in bridging technical implementation with business requirements and robust quality assurance.

Abstract

Generative AI systems promise transformative capabilities but present unique evaluation challenges due to their creative and unpredictable nature. Erin Pacquetet addresses this paradox by outlining comprehensive strategies for assessing systems that blend linguistic fluidity with strict factual accuracy. Using a Retrieval-Augmented Generation (RAG) chatbot as a running case study, the presentation examines limitations of traditional metrics, the role of LLM-as-a-judge approaches alongside their inherent biases, the necessity of human evaluation, and continuous monitoring to detect drift. Attendees gain practical frameworks for building reproducible evaluation pipelines that balance innovation with reliability in production environments.

The Fundamental Challenge of Evaluating Generative Systems

Generative AI introduces a core tension between creativity and control. Organizations adopt large language models precisely because they handle diverse, uncontrolled inputs and produce personalized outputs. Yet this very strength complicates evaluation. Traditional deterministic testing works for rule-based systems but falls short when outputs vary naturally while needing to remain accurate, relevant, and safe.

In the case study of an insurance company’s customer-facing RAG chatbot, the system must answer questions about policies while adhering to brand tone, regulatory constraints, and response length limits. A single question like “Is home insurance mandatory for tenants in France?” could yield multiple valid responses of varying quality. Evaluation must therefore move beyond binary correctness to nuanced assessment across multiple dimensions.

Effective evaluation pipelines transform qualitative judgments into quantitative, scalable measurements. This requires simulating realistic inputs, generating outputs, and assessing them against well-defined criteria. The process must cover ideal scenarios, expected real-world usage, and adversarial cases to ensure robustness before production deployment.

Simulating Inputs: Ideal, Realistic, and Adversarial Scenarios

The foundation of any evaluation lies in a carefully constructed dataset representing the full spectrum of potential interactions. For the insurance chatbot, inputs fall into three categories.

Ideal inputs are perfectly formed questions with clear intent and complete context, such as grammatically correct inquiries directly related to covered products. These establish baseline performance and set high acceptance thresholds.

Realistic inputs mirror actual user behavior: keyword-based queries, vague phrasing, oral-style language, partial context, or minor errors. Testing these ensures the system handles the messy reality of production traffic rather than sanitized examples.

Adversarial inputs probe vulnerabilities: prompt injections, attempts to elicit harmful content, off-topic questions, or malicious efforts to bypass safeguards. These reveal security weaknesses and edge cases that could damage reputation or expose risks.

Creating this dataset demands collaboration between technical teams and domain experts. Business stakeholders define what constitutes success for each category, translating abstract requirements into concrete examples. This exercise often reveals inconsistencies in initial specifications, forcing clarification before development advances.

The resulting evaluation dataset serves as both a benchmark and a living artifact. It evolves with the product, incorporating new failure modes discovered in production and expanding coverage as usage patterns emerge.

Generating and Assessing Outputs: Metrics and Human Judgment

Once inputs are prepared, the system generates outputs for evaluation. Assessment occurs along two primary axes: output quality and operational performance.

Output quality encompasses factual accuracy, relevance to the query, completeness of information, and safety. For the RAG chatbot, responses must draw correctly from policy documents, address the specific question asked, provide sufficient detail without excess length, and maintain an appropriate empathetic tone.

Traditional metrics prove insufficient. String matching fails to capture semantic equivalence across varied phrasings. Semantic similarity measures can overlook critical omissions or subtle inaccuracies. Probabilistic approaches, particularly LLM-as-a-judge, offer greater flexibility by leveraging models to analyze outputs against detailed criteria.

A well-crafted judge prompt might instruct the model to identify contradictions or omissions between a generated response and a reference answer, returning a binary judgment. This constrains the evaluation task sufficiently to reduce variance while maintaining nuance. Multiple specialized judges can target different aspects: one for factual consistency, another for tone alignment, and a third for regulatory compliance.

Human evaluation remains essential for validation. Domain experts review samples to calibrate automated metrics, ensuring alignment between machine judgments and business expectations. This human-in-the-loop process establishes confidence thresholds for each metric.

Operational metrics complement quality assessment. Response latency, cost per inference, and system stability must meet production requirements. A perfectly accurate but slow response fails as a product. Monitoring these dimensions alongside quality creates a holistic view of system readiness.

Building and Maintaining Evaluation Pipelines

A complete pipeline integrates input simulation, output generation, and multi-faceted assessment into an automated workflow. Teams execute evaluations frequently: after prompt modifications during development, before major releases, and continuously in production to detect regression or drift.

The evaluation dataset evolves as the central reference point. Production logs reveal new query patterns or failure modes, which teams incorporate to strengthen coverage. Regular human review sessions ensure metrics remain aligned with changing business needs and user expectations.

For the insurance chatbot, this meant balancing completeness against brevity, factual precision against approachable language, and safety against helpfulness. The dataset captured these trade-offs explicitly, allowing systematic optimization rather than guesswork.

Challenges persist. Judge models can inherit biases or exhibit inconsistency. Human evaluators introduce subjectivity. Thresholds require careful tuning to avoid both false confidence and excessive caution. Success demands iterative refinement and cross-functional collaboration.

From Evaluation to Production Confidence

Robust evaluation bridges the gap between promising prototypes and reliable production systems. By systematically addressing the inherent variability of generative outputs, teams build the confidence necessary for deployment.

The insurance chatbot case demonstrates that evaluation is not merely technical validation but a strategic discipline. It forces clarification of requirements, surfaces hidden assumptions, and creates shared understanding across technical and business stakeholders.

As generative AI proliferates, organizations that master evaluation gain competitive advantage. They deploy innovative capabilities with appropriate safeguards, iterating rapidly while maintaining quality. The discipline transforms the “unmeasurable” into something manageable, turning potential risk into sustainable value.

Links:

PostHeaderIcon [AWSReInvent2025] Breaking Performance and Cost Barriers in Generative AI: The Strategic Role of AWS Trainium

Lecturer

Gadi Hutt is a Senior Director of Product Management at AWS, specializing in the development and strategic scaling of specialized silicon. With an extensive background in semiconductor engineering and cloud infrastructure, Gadi has been a pivotal figure in the evolution of the AWS Annapurna Labs team. His work focuses on delivering high-performance, cost-efficient compute solutions that address the exponential resource demands of modern artificial intelligence. He is joined by industry leaders such as Joe Spisak, Product Director at Meta, and Oren Shomar, Director of Engineering at poolside, who provide empirical evidence of the impact of these specialized chips on global AI model development.

Abstract

The rapid proliferation of generative artificial intelligence (GenAI) has introduced unprecedented computational challenges, characterized by skyrocketing training costs and intricate scaling requirements. This article examines the architectural innovations of AWS Trainium2, the second-generation purpose-built chip designed specifically for high-performance deep learning. By analyzing the integration of Trainium2 into the AWS UltraCluster environment and the supporting Neuron SDK, we explore how specialized silicon provides a viable alternative to general-purpose GPUs. The discussion highlights real-world applications by Meta and poolside, demonstrating significant gains in price-performance for training Mixture of Experts (MoE) models and deploying agentic systems. Furthermore, the article outlines the methodological shift toward optimized software-hardware co-design as a necessity for sustaining the next generation of AI innovation.

The Architectural Foundation of Purpose-Built Silicon

The foundational shift in AI infrastructure is driven by the realization that general-purpose hardware often encounters bottlenecks when processing the massive parameter counts of modern Large Language Models (LLMs). Gadi explains that AWS Trainium2 was engineered to alleviate these constraints by focusing on three primary pillars: compute density, high-speed interconnectivity, and memory efficiency.

A critical innovation in this generation is the transition to a more robust node technology that allows for significantly higher teraflops (TFLOPS) per chip compared to its predecessor. This is complemented by the AWS Nitro System, which offloads networking and storage functions, allowing the Trainium processors to dedicate nearly 100% of their resources to model arithmetic. The architecture supports a diverse range of data types, including FP8 and Transformer Engine optimizations, which are essential for maintaining precision while reducing computational overhead.

Scaling with AWS UltraClusters and Elastic Fabric Adapter

Individual chip performance is only one aspect of the solution; the ability to scale to tens of thousands of chips is where the true breakthrough occurs. Gadi describes the AWS UltraCluster as a massive, non-blocking network of Trainium2 instances connected via the second-generation Elastic Fabric Adapter (EFA). This infrastructure enables petabit-scale networking, which is crucial for the frequent synchronization required during distributed training.

The EFA technology utilizes a custom-built protocol designed to minimize latency and jitter, which are often the limiting factors in synchronous training workloads. By providing a high-bandwidth, low-latency fabric, AWS allows developers to treat an entire cluster of thousands of nodes as a single, unified computer. This capability is particularly relevant for training foundational models where the dataset and model weights are too large to fit into the memory of a single machine.

Industry Validation: Meta and the Llama Ecosystem

The practical utility of Trainium2 is underscored by its adoption by major industry players. Joe Spisak from Meta highlights the collaborative effort to integrate Trainium2 into the Llama model ecosystem. For a company operating at Meta’s scale, the primary objective is to maximize “tokens per dollar.”

Joe notes that the integration of Trainium2 with the PyTorch framework via the AWS Neuron SDK allows Meta to leverage their existing codebases while benefiting from the superior price-performance of AWS silicon. This partnership demonstrates that purpose-built hardware can successfully support the most demanding open-source model architectures, providing the global community with more efficient paths to fine-tuning and deploying sophisticated AI systems.

Case Study: High-Efficiency Training at poolside

Oren Shomar from poolside provides a deep dive into the specific challenges of building AI for software engineering. Their workload requires massive-scale training on code repositories, which involves long-sequence lengths and complex reasoning patterns. poolside transitioned to Trainium2 to overcome the cost barriers associated with traditional GPU clusters.

Oren emphasizes the role of the Neuron SDK in this transition. The compiler’s ability to automatically optimize graph execution and manage memory across the Trainium cores was a decisive factor in achieving their performance targets. By using Trainium2, poolside was able to maintain a rapid iteration cycle, training new model variants in a fraction of the time and cost previously required, thereby accelerating their path to delivering agentic reasoning capabilities to developers.

The Neuron SDK: Bridging Frameworks and Silicon

The success of specialized silicon is inextricably linked to the software stack that exposes its power. The AWS Neuron SDK acts as the interface between popular machine learning frameworks like PyTorch and JAX and the underlying Trainium hardware.

The Neuron compiler performs sophisticated optimizations, including operator fusion and tensor tiling, to ensure that the hardware is utilized at peak efficiency. Gadi highlights the “Neuron Distributed” library, which provides high-level abstractions for data parallelism, pipeline parallelism, and tensor parallelism. This allows researchers to scale their models across an UltraCluster without having to manually manage the complexities of collective communication or device-specific memory management.

Conclusion: The Imminent Future of AI Infrastructure

The trajectory of GenAI necessitates a departure from the “one-size-fits-all” hardware approach. Through the development of Trainium2 and the accompanying ecosystem, AWS has established a new benchmark for scalable AI training. Gadi concludes that the commitment to continuous innovation—evidenced by the early announcement of Trainium4—ensures that the industry can keep pace with the evolving complexity of AI models. As price-performance becomes the dominant metric for AI viability, specialized silicon like Trainium will be the cornerstone of a sustainable and innovative technological future.

Links:

PostHeaderIcon [AWSReInvent2025] Accelerating E-Commerce Insights with Snowflake Intelligence: A Case Study on Decile’s Luma AI Analyst

Lecturer

Santiago Giraldo serves as Senior Director of Product Marketing for Artificial Intelligence at Snowflake. With over 15 years of experience in data and AI technology, Santiago specializes in bridging business needs with advanced technical solutions, focusing on generative AI and enterprise data platforms. He holds a background from Parsons School of Design – The New School and is based in the Denver Metropolitan Area.

Brian Neumann is Senior Vice President of Engineering at Decile, an e-commerce analytics platform. Brian leads engineering efforts to develop innovative data solutions for brands, emphasizing multi-tenant architectures and AI integration to enhance customer insights.

Abstract

This presentation explores the transformative potential of Snowflake Intelligence, a generative AI-powered feature set designed to enable natural language interactions with enterprise data. Santiago introduces the core principles of Snowflake Intelligence, addressing longstanding challenges in data accessibility and decision-making velocity. Brian then details Decile’s implementation, showcasing how the platform powers Luma, a custom AI analyst that democratizes e-commerce insights across organizational roles. The discussion highlights architectural strategies, trust mechanisms, and practical outcomes, illustrating how agentic AI can shift enterprises from reactive reporting to proactive, reasoned action.

Bridging the Gap Between Business and Data Teams

Enterprises often grapple with disparities in how business users and data teams interact with information. Business stakeholders require timely, actionable insights to drive decisions, yet data teams frequently dedicate substantial effort to producing static reports or dashboards. By the time these deliverables reach decision-makers, opportunities may have diminished, as insights arrive too late for effective intervention.

Snowflake Intelligence addresses this divide by empowering users—from executives to frontline employees—to pose complex questions in natural language and receive reasoned responses. Unlike traditional tools limited to surface-level queries (e.g., “What were sales last week?”), this innovation facilitates deeper inquiry, such as identifying underlying causes or forecasting future trends. It integrates data from disparate sources, including databases, customer platforms like Salesforce, and third-party enrichments, all within a secure, governed environment.

A key advantage lies in its enterprise readiness: features are native to the Snowflake platform, ensuring robust governance, security, and data quality. This approach fosters a “reasoning partner” dynamic, where AI not only retrieves data but also provides explanatory context, enabling high-confidence decisions in real time.

Core Principles and Capabilities of Snowflake Intelligence

Snowflake Intelligence rests on three foundational pillars: deep analysis, trust, and enterprise-grade security.

Deep analysis extends beyond descriptive reporting to prescriptive and predictive reasoning. Users can explore questions like “What headwinds threaten upcoming sales?” or “How can retention be improved?” by leveraging multimodal data—structured and unstructured—across the organization. Features such as research mode enable forward-looking investigations, drawing from comprehensive knowledge sources.

Trust is paramount in generative AI adoption, where hallucinations or opaque reasoning erode confidence. Snowflake mitigates this through verified answers, full traceability to original sources, and transparent explanations. Responses include reformulated queries, step-by-step reasoning, and direct links to underlying SQL, allowing verification down to individual data points.

Enterprise readiness ensures all operations occur within Snowflake’s governed ecosystem. Dynamic discovery provides clear explanations of results, while integrations with marketplace data and enterprise tools unify insights. This holistic design transforms data utilization, placing organizational knowledge at users’ fingertips for instantaneous, reliable exploration.

Decile’s Journey: From Traditional Analytics to AI-Driven Insights

Decile operates as an e-commerce analytics platform, serving over 100 leading brands by aggregating data from sources like Shopify, Magento, marketing channels, and enrichment providers such as Acxiom. The platform creates dedicated Snowflake data warehouses per client, overlaid with application experiences for lifecycle reporting and customer segmentation.

Initially, Decile’s dashboards aimed to surpass native platform reporting by stitching disparate data for richer views. However, brand variability—ranging from retail integrations to diverse product analytics—complicated dashboard flexibility. This led to increased complexity for non-technical users, who grew reliant on customer success teams, effectively positioning Decile as an outsourced data function.

Recognizing this barrier, Decile sought to empower clients directly through an AI analyst. Early prototyping with various frameworks revealed significant hurdles: building vector stores for semantic understanding, ensuring SQL accuracy, providing visualizations, and establishing evaluation mechanisms. These challenges posed substantial investment risks for a startup.

Snowflake Intelligence emerged as an ideal solution, leveraging existing governed warehouses and dbt semantic models. Implementation involved extending dbt documentation with metadata (aliases, synonyms, sample values) to inform semantic views—YAML-defined structures describing dimensions, measures, relationships, and natural language descriptions.

Cortex Search services enhanced fuzzy matching for free-form queries, while verified queries predefined complex calculations (e.g., retention cohorts). A custom library automated provisioning of semantic views, search services, and agents via deployment pipelines.

Implementation Outcomes and Future Directions at Decile

Rapid deployment enabled pilot testing through Snowflake’s UI, gathering feedback to refine models. API access facilitated seamless integration into Decile’s application, branding the experience as Luma—a conversational AI analyst.

Users reported substantial time savings, with marketers and executives conducting analyses previously requiring extensive report stitching. Visible thinking steps—detailing semantic mappings and reasoning—built confidence, reducing perceived black-box risks. Support queries dropped 75% among adopters, as users self-served segments for activation.

Luma’s impact extends to operational efficiency: quicker market entry, reduced custom report demands, and empowered segmentation (e.g., identifying repeat purchasers for subscriptions).

Looking ahead, Decile plans per-brand instruction customization to capture nuances (e.g., subscription vendors, wholesale handling). Aspirations include user-contributed context for vertical-specific analyses and scheduled alerting for anomalies, emulating a proactive human analyst.

Implications for Enterprise AI Adoption

This collaboration exemplifies how Snowflake Intelligence lowers barriers to agentic AI in specialized domains. By providing turnkey frameworks—semantic views, APIs, and observability—platforms like Decile accelerate innovation without prohibitive development overhead.

Broader implications include democratized data access, reducing silos and delays while upholding trust through traceability. For e-commerce, this translates to agile responses to market dynamics, personalized strategies, and sustained growth.

Ultimately, such integrations signal a shift toward AI-augmented workflows, where tools complement human expertise, fostering cultures of data-driven agility and innovation.

Links:

PostHeaderIcon [PyDataGlobal2025] Using Traditional AI and Large Language Models to Automate Complex and Critical Documents in Healthcare

Lecturer

Lily Xu is a Data Science Director in the corporate data-science and AI team at Vertex Pharmaceuticals, where she has worked for approximately seven years. She leads interdisciplinary groups of data scientists, data engineers, software engineers, and operations specialists focused on clinical-area solutions. She holds a doctorate in bioengineering from the Massachusetts Institute of Technology and an undergraduate degree from the University of California, Berkeley. Her earlier research produced publications on virtual microfluidics and the human microbiome; at Vertex she has driven projects spanning generative AI for clinical documentation, predictive patient modeling, large-scale claims analytics, protocol design, and centralized site intelligence.

Abstract

Informed consent forms constitute high-stakes, patient-facing, heavily regulated documents that must be tailored to jurisdictional requirements, local ethics boards, and plain-language standards. Their manual production across dozens of countries and hundreds of sites creates substantial operational bottlenecks in clinical-trial start-up. This article examines a production system developed at Vertex Pharmaceuticals that combines classical document-processing pipelines with large language models to auto-draft informed consent forms at scale. Emphasis is placed on architectural choices that minimize hallucination risk, rigorous measurement of end-to-end time savings, the centrality of change management, and the longer-term strategy of constructing a connected document network rather than isolated point solutions.

Clinical-Trial Operations Context and the Dual AI Portfolio

Clinical-trial operations span design, planning, execution, and monitoring phases, each generating or consuming large volumes of structured and unstructured documents. Failure to recruit patients, suboptimal site selection, or protracted regulatory review can each cost tens to hundreds of millions of dollars. Beginning in 2019 the Vertex data-strategy and solutions team—functioning as an internal SWAT unit—built trust through small, measurable pilots that combined public and private data into AI-ready assets. Over successive years the portfolio matured from ad-hoc analytics into standardized offerings for site identification, patient finding, enrollment forecasting, and, more recently, generative document automation.

The team deliberately distinguishes analytical AI (predictive modeling, Bayesian enrollment forecasts, rare-disease patient identification) from generative AI (first-draft document creation, knowledge-base chat, brand-copy generation). Business partners often approach the group believing a problem requires generative technology when structured data and classical machine learning would suffice; conversely, generative methods unlock previously intractable free-text tasks. Framing the two categories helps both data scientists and operational stakeholders select the appropriate tool. A foundational data layer aggregates site performance metrics, physician databases, claims, and census information; disease-specific analytic views and predictive models sit atop this foundation. Parallel generative pipelines extract structured content from lengthy protocols and feed downstream document generators, with embedded quality-control workflows so that extraction errors are corrected before they propagate into patient-facing material.

Architecture of the Informed-Consent-Form Auto-Drafting System

An informed consent form must convey risks, procedures, and rights in plain language while satisfying country-specific and sometimes site-specific regulatory requirements. A single multi-country trial may therefore require dozens of distinct variants. The solution developed at Vertex treats the clinical protocol as the primary source of truth, a blank regulatory template as the structural skeleton, and an approved language library as the repository of standardized phrasing.

Custom Python modules parse the protocol into logically coherent sections rather than arbitrary token chunks. Section-specific prompts and deterministic extraction routines pull the necessary facts. User-supplied answers to questions that cannot be parsed from the protocol are collected through a controlled interface. The resulting structured payload is inserted into the template; approved language snippets are retrieved via API from a purpose-built library that replaced earlier Excel spreadsheets and now maintains full audit trails and disease-area tagging.

The application is implemented in Flask and Dash, hosted on AWS behind single-sign-on, and calls a private Microsoft OpenAI endpoint for the generative steps. A monitoring dashboard continuously compares newly generated drafts against ground-truth forms produced by human experts, allowing the team to detect drift in accuracy over time. Because the generative component constitutes only a minority of the code base, the majority of engineering effort is devoted to robust parsing, template management, and workflow orchestration—skills that remain essential even as language models improve.

The design philosophy is “AI in the human loop” rather than “human in the AI loop.” Regulatory and patient-safety constraints demand that every draft undergo expert review; the system’s value lies in accelerating the initial drafting phase so that reviewers begin from a high-quality baseline rather than a blank page.

Measuring Impact, Change Management, and Scaling Strategy

Early controlled experiments compared pure manual drafting (one to three hours depending on trial complexity) with auto-draft generation (under ten minutes). Drafting-time reduction approached 90 percent. When subsequent editing and quality-control effort was included, net end-to-end time savings settled near 40 percent—still substantial given the volume of forms required across a growing portfolio. Because operational teams are chronically time-constrained, such measurements were performed on only two trials; the results nevertheless provided the quantitative foundation for continued investment.

Technology alone does not guarantee adoption. Change-management activities therefore received equal attention: standardization of templates and language libraries, transparent communication of model assumptions and known failure modes, and staged training that enabled business users to generate drafts independently. Treating free-text language assets with the same governance rigor applied to numerical data proved essential.

The longer-term vision is a connected document network rather than a collection of isolated point solutions. Clinical protocols and clinical study reports function as central hubs; mapping the full input–output relationships among start-up documents reveals opportunities for shared extraction components and cascading automation. The same platform is being extended to site budgets, case-report-form specifications, training materials, and other protocol-derived artifacts. Country-level templates are already linked so that a single protocol upload can spawn multiple jurisdiction-specific drafts simultaneously. Site-level customization remains outside the automated scope because the return on investment diminishes rapidly at that granularity; country-level guidance is instead provided to local teams.

Broader Lessons for Generative Applications in Regulated Environments

Several observations travel beyond the specific use case. First, impact measurement must be designed from the outset; without side-by-side timing studies and accuracy tracking, claims of productivity gain remain anecdotal. Second, the majority of engineering effort in production document systems continues to reside in classical software and data-engineering practices; large language models occupy a focused niche once reliable extraction and templating are in place. Third, alignment with business ownership is decisive: projects lacking motivated operational sponsors are deferred in favor of those with clear accountability and enthusiasm. Finally, the cumulative benefit of a systematically constructed document network can outweigh the initial development cost provided the organization persists past the early pilots.

Ambient listening, internal retrieval-augmented generation over institutional knowledge bases, and protocol optimization via real-world data are complementary initiatives already underway at Vertex and peer organizations. Collectively they illustrate a measured trajectory in which generative and analytical methods remove routine cognitive load while leaving critical reasoning and final accountability with domain experts.

Links:

PostHeaderIcon [MiamiJUG] Retrieval-Augmented Generation: Building Deterministic AI for Production

Lecturer

Frank Greco is a Java Champion, enterprise architect, and senior consultant specializing in Artificial Intelligence and Cloud computing. He is the founder and Chairman of NYJavaSIG and a co-author of JSR #381 “VisRec,” the Java API for visual recognition. Frank is a recognized educator and technical leader who has presented at major global conferences including JavaOne, DevNexus, and Devoxx.

Abstract

This article provides an analytical framework for integrating Large Language Models (LLMs) into production Java environments using Retrieval-Augmented Generation (RAG). By moving beyond simple chat interfaces to programmatic API access, developers can build AI systems that are grounded in verified enterprise data. The analysis explores prompt engineering methodologies—such as Few-Shot and Chain of Thought (CoT)—and the architectural role of vector databases in mitigating model hallucinations while ensuring data security and version control.

Methodologies in Prompt Engineering

Prompting is the primary mechanism for steering the behavior of a neural network. Unlike traditional programming, prompting is probabilistic rather than deterministic. Frank identifies several advanced techniques to improve model reliability:

  • Zero-Shot and Few-Shot Learning: Few-shot prompting provides the model with specific examples of the desired input-output pattern, significantly improving the accuracy of complex tasks.
  • Chain of Thought (CoT): This instructs the model to “think step-by-step,” detailing its reasoning process before providing a final answer. This methodology is critical for reducing logical errors.
  • Persona Identification: Assigning a specific role to the model (e.g., “Act as a Java security expert”) helps contextualize the response and refine the output tone.

Architectural Implementation: Retrieval-Augmented Generation (RAG)

To overcome the limitations of an LLM’s static training data, enterprises utilize RAG to ground the model in real-time, private data. In a RAG architecture, a user query is first used to search a knowledge base—typically a Vector Database—for relevant documents. This retrieved context is then injected into the prompt, allowing the LLM to generate an answer based on specific facts rather than general probabilities.

This approach offers several production-grade benefits:

  1. Reduced Hallucinations: By providing the model with the necessary facts, the likelihood of it “making up” information is significantly decreased.
  2. Data Security: RAG allows models to use private company information without that data being used to train the underlying public model.
  3. Traceability: Responses can be cited back to specific source documents found in the vector database.

Production Challenges and Ethical Considerations

Implementing AI at scale introduces significant engineering overhead. Developers must manage Prompt Versioning to ensure consistent behavior across deployments and navigate the legal implications of AI-generated content. Furthermore, because these are probabilistic systems, Frank warns that if a wrong answer poses a high risk to the business, generative AI may not be the appropriate solution. Engineers must balance the productivity gains of AI with the need for rigorous safety guardrails and human-in-the-loop verification.

Links:

PostHeaderIcon [PyDataGlobal2025] What’s Next in AI for Data and Data Management

Lecturer

Lisa Amini is a Distinguished Engineer at IBM and Director of Data & AI Platforms Research, where she also leads IBM’s AI Horizons Network. Her career at IBM Research spans more than two decades and includes foundational work on stream processing systems that became the InfoSphere Streams product, leadership of the IBM Research laboratory in Ireland, and earlier roles directing knowledge and reasoning research. She has guided interdisciplinary efforts across cloud computing, artificial intelligence, and quantum computing, always with an emphasis on technologies that can be deployed at enterprise scale.

Abstract

Recent advances in large language models have catalyzed a wave of AI-assisted tools for data management and operations, ranging from code-generation assistants for data-flow pipelines to retrieval-augmented generation systems and increasingly autonomous data agents. This keynote examines the rapid evolution of generative and agentic capabilities, situates them within the broader data-management stack, and explores both near-term practical applications and longer-horizon research directions. Particular attention is given to the shift from human-operated systems augmented by copilots toward semi-autonomous stacks in which agents design, optimize, remediate, and continuously evaluate data products. The discussion balances technical opportunity with the enduring requirements of price-performance, open-source interoperability, and hybrid data architectures.

The Accelerating Capability Curve and the Emergence of Agency

The pace at which machine-learning benchmarks reach human-level performance has changed dramatically. Tasks that once required decades of incremental progress—handwriting recognition, for example—now reach parity within a few years or even months. Reading comprehension and predictive reasoning benchmarks follow similarly steep trajectories. While these evaluations remain narrow and do not constitute artificial general intelligence, they illustrate an unprecedented rate of improvement. Simultaneously, the cost per inference continues to fall even as model size and training compute grow, a trend driven by better systems design and algorithmic efficiency.

Within this landscape the progression from predictive models to generative models to conversational systems and finally to agents marks a qualitative shift. Agents do not merely answer questions; they dynamically control application flow, make decisions, take actions, and attempt self-correction. In the data domain this agency opens the possibility of systems that no longer wait for humans to formulate every query or repair every broken pipeline. Instead, agents can probe schema, resolve ambiguity, hypothesize data products, evaluate their own output, and iterate.

Transforming the Data Landscape and the Complementary Task Stack

Unstructured data has long existed, yet only recently has it assumed central importance. Machines can now reason over images, generate multimodal content, and extract structured signals from free text at scale. Classical database, warehouse, and lakehouse architectures, optimized primarily for structured tables, must therefore accommodate new access patterns. Retrieval-augmented generation pipelines replace static queries with dynamic retrieval-plus-generation cycles. User interaction moves from fixed application-generated SQL toward speculative, multi-step agent dialogues that probe metadata, formulate candidate queries, and refine them in light of intermediate results.

A useful conceptual inversion is to view the traditional storage–compute–query stack alongside a complementary human-task stack: infrastructure design, workload optimization, data discovery, enrichment, flow creation, remediation, governance, and insight generation. Each of these human activities constitutes fertile ground for agentic automation. Early systems already demonstrate learnable components inside query optimizers and routers; more ambitious research explores whether agents can search the design space of kernel-level software itself.

From Automation to Autonomy: Data Products and Continuous Evaluation

The practical goal is not merely to accelerate individual steps but to move entire workflows from human-operated to human-supervised. Consider the request to stand up a data stack and associated data products for a new application—robo-trading, for instance, that must combine public market data with sentiment signals and support periodic rebalancing. A multi-agent system can be tasked with discovering relevant sources, hypothesizing an ideal schema, populating that schema from heterogeneous tables and documents, extracting structured fields from natural-language text, and packaging the result as a governed data product.

Critical to autonomy is the ability to evaluate quality without constant human intervention. One effective strategy generates natural-language questions that a domain expert would plausibly ask of the intended data product, translates those questions into executable queries, and then monitors coverage metrics (tables and columns touched), topic coverage, query complexity, and latency. Agents iterate—adding sources, refining transformations, simplifying views—until the metrics stabilize within acceptable bounds or progress plateaus and human guidance is required. The same loop can later serve as continuous monitoring: questions that once succeeded can be re-executed to detect drift or regression.

Similar patterns apply to operational remediation. When a data-flow pipeline fails, agents can examine logs, generate natural-language root-cause hypotheses, propose script repairs, and, under appropriate guardrails, test those repairs in a sandbox before presenting them for approval. Across the spectrum of use, build, and optimize activities, the user’s role gradually shifts from operator to approver or observer.

Enduring Constraints and the Research Horizon

Price-performance remains non-negotiable; open-source components continue to enable rapid composition of storage formats, query engines, and table formats; hybrid architectures that span on-premises, cloud, and edge locations persist. Benchmarks, data contracts, open lineage standards, and carefully scoped open-weight models supply the interfaces and evaluation harnesses that allow agents to interoperate safely. Research prototypes already explore operator libraries that let developers request high-level transformations while large language models synthesize the concrete implementations behind the scenes.

The path forward is incremental. Fully autonomous data stacks will not appear overnight. Yet the combination of generative models, agent frameworks, and rigorous evaluation loops is already moving concrete workloads—data-product curation, flow repair, insight generation—along the continuum from assistance toward autonomy. The opportunity for data scientists and engineers is to shape the metrics, tools, and governance practices that will keep these systems both powerful and trustworthy.

Links:

PostHeaderIcon [AWSReInvent2025] Transforming Integrated Diagnostics: Philips’ AI-Driven Evolution on AWS

Lecturer

Sam Cool is a Director and Global Lead for Healthcare Solutions at Amazon Web Services (AWS), where he focuses on accelerating digital transformation for global health organizations. With extensive experience in cloud architecture and clinical workflows, Sam works with industry leaders to dismantle data silos and implement scalable AI solutions. Jared Nicks is a Principal Solutions Architect at AWS, specializing in medical imaging and Health-IT. His work is instrumental in developing the AWS HealthImaging service, which provides high-performance storage and retrieval for large-scale medical datasets. Wilson Toe serves as a Senior Product Manager at AWS, focusing on the intersection of Generative AI and healthcare analytics. Dr. Praeloski is a Senior Clinical Scientist at Philips, bringing decades of expertise in diagnostic imaging, pathology, and cardiology. He leads Philips’ efforts to integrate multi-modal data into a unified platform that enhances clinical decision-making. Together, these experts have pioneered a collaboration that leverages cloud-native technologies to redefine the diagnostic landscape.

Abstract

Modern healthcare is characterized by an explosion of diagnostic data, yet this information remains largely fragmented across disparate systems for radiology, cardiology, and pathology. This fragmentation hampers the ability of clinicians to form a holistic view of the patient, leading to diagnostic delays and suboptimal treatment planning. This article examines the strategic journey of Philips in transforming integrated diagnostics through its partnership with AWS. By shifting from on-premises infrastructure to a cloud-native architecture, Philips has successfully integrated diverse data streams, with a particular focus on the emerging frontier of digital pathology. The discussion explores the technical implementation of AWS HealthImaging, the transition to standardized DICOM formats for pathology, and the application of Generative AI to streamline clinical reporting. Ultimately, this framework enables global collaboration and real-time diagnostic consensus, moving the needle toward truly personalized and precise medicine.

The Paradox of Fragmented Diagnostic Intelligence

The clinical diagnostic process is the cornerstone of patient care, influencing over 70% of medical decisions. However, the current infrastructure supporting these decisions is often a patchwork of “black boxes.” A patient’s journey typically involves multiple diagnostic touchpoints: an X-ray in radiology, an ECG in cardiology, and a tissue biopsy in pathology. Historically, each of these domains has operated in a silo, utilizing proprietary data formats and isolated storage systems. Sam observes that while the volume of data is increasing—driven by higher-resolution imaging and molecular diagnostics—the “intelligence” derived from that data remains localized.

For a clinician, this fragmentation means navigating multiple interfaces and manually correlating reports, a process prone to error and inefficiency. The transition to integrated diagnostics is not merely a technical upgrade; it is a clinical necessity. By centralizing these streams in the cloud, healthcare providers can move from a reactive, department-centric model to a proactive, patient-centric one. Philips’ vision for integrated diagnostics centers on breaking down these silos to provide a “single source of truth” for every patient, regardless of where the data was generated.

Digital Pathology: The Final Frontier of Digitalization

While radiology and cardiology have been digital for decades, pathology—the study of tissue samples—has remained stubbornly analog. For over a century, pathologists have relied on glass slides and manual microscopy. The sheer scale of the data involved has been the primary barrier; a single high-resolution digital slide can exceed several gigabytes in size, and a single patient case may involve dozens of slides.

Dr. Praeloski highlights that digital pathology represents the next great shift in clinical innovation. By digitizing these slides, Philips enables pathologists to work in an environment that is “born digital,” allowing for the application of computer vision and machine learning. This transition is facilitated by the adoption of the DICOM (Digital Imaging and Communications in Medicine) standard for pathology images. Standardizing these massive datasets allows them to be treated with the same rigor and interoperability as traditional radiological images, enabling them to be stored, shared, and analyzed within the same AWS-backed ecosystem.

Architecting for High-Throughput Imaging with AWS HealthImaging

The technical challenge of managing millions of high-resolution pathology slides requires an infrastructure that can handle extreme throughput and low-latency retrieval. Standard object storage, while durable, often struggles with the specific access patterns required for medical imaging, where a clinician needs to “zoom and pan” through a multi-gigabyte image in real-time.

To solve this, Philips leverages AWS HealthImaging. This purpose-built service allows for the ingestion of medical images at scale while providing sub-second access to specific image frames. By decoupling storage from the viewing application, AWS HealthImaging ensures that clinicians can access images from any device, anywhere in the world, without the need for high-powered local workstations.

'''# Conceptual example of fetching metadata for a DICOM image set'''
import boto3

health_imaging = boto3.client('healthimaging')

def get_image_metadata(datastore_id, image_set_id):
    response = health_imaging.get_image_set_metadata(
        datastoreId=datastore_id,
        imageSetId=image_set_id
    )
    return response['metadata']

Jared emphasizes that this architecture is foundational for “high-throughput” clinical environments. In a traditional setup, moving a slide from storage to a viewer could take minutes; with HealthImaging, it takes milliseconds. This efficiency is critical in pathology, where time-to-diagnosis directly impacts patient outcomes in oncology and acute care.

Empowering Clinicians through Generative AI and Automated Reporting

Once diagnostic data is centralized and accessible, the next challenge is synthesis. Pathologists and radiologists spend a significant portion of their day dictating and transcribing findings. Generative AI offers a transformative solution by automating the creation of structured reports and summarizing complex longitudinal patient histories.

Wilson explains how Philips integrates Amazon Bedrock to assist in the “last mile” of the diagnostic process. By analyzing the metadata and AI-detected features of an image, the system can draft a preliminary report that the clinician then reviews and validates. This doesn’t replace the expert; rather, it removes the “blank page” problem and ensures that reports follow a standardized, high-quality format. Furthermore, LLMs (Large Language Models) can scan years of a patient’s prior records to highlight relevant changes—such as the growth of a lesion over time—that might be missed in a manual review.

Global Collaboration and the Future of Consensus

One of the most profound impacts of shifting integrated diagnostics to the cloud is the enablement of global collaboration. In the analog world, seeking a second opinion on a rare pathology case required physically shipping glass slides across borders—a process that was slow, expensive, and risky.

Through Philips’ cloud-native platform, a specialist in New York can consult on a case in London in real-time. The digital platform supports “shared view” sessions where multiple clinicians can annotate the same slide simultaneously. Dr. Praeloski notes that in recent surveys, 100% of pathologists using the digital system reported that it facilitated reaching a diagnostic consensus more effectively than manual methods. This democratization of expertise is particularly vital for underserved regions, where access to specialized sub-pathologists is limited.

Conclusion: A Paradigm Shift in Precision Medicine

The journey of Philips and AWS illustrates that the future of healthcare is not just about “better machines,” but about “smarter data.” By integrating radiology, cardiology, and pathology into a unified cloud-native framework, they have laid the groundwork for the next generation of precision medicine. This evolution reduces clinical burnout by automating administrative tasks, improves diagnostic accuracy through AI assistance, and accelerates the pace of care through global collaboration. As the system continues to scale, the data captured today will become the training ground for the cures of tomorrow, proving that when diagnostic intelligence is integrated, the potential for clinical innovation is limitless.

Links:

PostHeaderIcon [AWSReInvent2025] Agentic AIOps: Navigating the Paradigm Shift toward Autonomous IT Operations

Lecturer

Abhijit Chakravarty, Mike Bechtel, and Michael J. Kavis
Abhijit Chakravarty is a seasoned technology leader at LogicMonitor, focusing on the intersection of artificial intelligence and infrastructure monitoring. Mike Bechtel serves as the Chief Futurist at Deloitte Consulting LLP, where he leads research into emerging technologies and their long-term impact on the enterprise. Michael J. Kavis is a Managing Director at Deloitte Consulting and a renowned expert in cloud computing and enterprise architecture, having authored multiple books on cloud transformation. Together, they represent a convergence of industry-leading monitoring solutions and strategic advisory expertise, specifically targeted at preparing global organizations for the complexities of the agentic AI era.

Abstract

As enterprise IT environments grow in scale and complexity, traditional AIOps frameworks—which primarily focused on pattern recognition and anomaly detection—are evolving into “Agentic AIOps.” This article explores the conceptual transition from systems that merely observe and alert to autonomous agents capable of reasoning, planning, and executing remediation tasks. By examining the integration of Large Language Models (LLMs) with operational telemetry, the study highlights a methodology centered on reducing “mean time to repair” (MTTR) and minimizing human intervention in repetitive incident management cycles. The analysis delves into the maturity model for agentic adoption, the necessity of rigorous data grounding, and the evolving role of the human operator in a supervised autonomous ecosystem. The findings suggest that agentic AIOps is not merely an efficiency tool but a fundamental redesign of IT governance and service reliability.

The Conceptual Evolution: From Observability to Autonomy

The IT landscape has historically progressed through distinct phases of monitoring. Early systems were reactive, relying on static thresholds to trigger alerts. This gave way to the first generation of AIOps, which utilized machine learning for event correlation and root cause analysis. However, even these advanced systems remained largely “human-in-the-loop,” where the AI identified a problem, but a person had to decide and act on the solution.

Agentic AIOps represents a paradigm shift where the AI moves from an advisor to a doer. Unlike traditional automation, which follows a rigid, pre-defined script (e.g., “if X, then do Y”), agentic systems utilize the reasoning capabilities of LLMs to handle “non-deterministic” scenarios. These agents can interpret natural language incident reports, query multiple databases to gather context, and generate a step-by-step remediation plan that adapts to the specific nuances of the failure.

Methodology: Reasoning, Tool-Use, and Grounding

The architecture of a modern agentic AIOps system, such as LogicMonitor’s “Edwin AI,” relies on three core pillars: reasoning, tool-use, and grounding.

Strategic Reasoning and Planning

The “brain” of the agent is the LLM, which processes incoming alerts through a reasoning framework—often employing the “ReAct” (Reason + Act) pattern. When an incident occurs, the agent first decomposes the problem into smaller, manageable sub-tasks. It formulates a hypothesis about the root cause and identifies the necessary information required to validate that hypothesis.

Dynamic Tool-Use

To act on its reasoning, the agent must be able to interact with the environment. This is achieved through “function calling” or tool-integration. An agent might have access to a suite of tools, including:

  • Infrastructure APIs: To restart services, scale resources, or modify configurations.
  • Knowledge Bases: To retrieve historical documentation or runbooks.
  • Communication Platforms: To update Slack channels or create ServiceNow tickets.

The Grounding Requirement

A critical challenge in applying generative AI to IT operations is “hallucination.” To ensure the agent makes decisions based on facts rather than probability, the methodology emphasizes “grounding” via Retrieval-Augmented Generation (RAG). The system feeds the LLM real-time telemetry from LogicMonitor alongside enterprise-specific runbooks. This ensures that the agent’s reasoning is constrained by the actual state of the infrastructure and the organization’s approved operating procedures.

Implementation: The Agentic Maturity Model

Adopting agentic AIOps is not an “all-or-nothing” proposition; it follows a maturity curve that balances autonomy with risk management.

  1. Assisted Mode: The agent acts as a co-pilot, summarizing incidents and suggesting remediation steps to a human operator for approval.
  2. Supervised Autonomy: The agent executes low-risk tasks autonomously (e.g., clearing disk space) while requiring permission for higher-impact changes (e.g., rebooting a production database).
  3. Full Autonomy: The system operates independently within strictly defined guardrails, only involving humans for unprecedented or catastrophic failures.

This tiered approach allows organizations to build trust in the agent’s decision-making while gradually reducing the cognitive load on Site Reliability Engineering (SRE) teams.

Consequences for Enterprise IT and the Workforce

The shift toward agentic operations necessitates a change in the mindset of IT leadership. The focus moves from “managing tasks” to “managing outcomes.” The role of the human operator evolves from a manual troubleshooter to a “curator of intent.” Engineers will spend less time reacting to pagers and more time defining the policies, objectives, and constraints within which the agents must operate.

Furthermore, the integration of LogicMonitor with platforms like Worldwide Technologies (WWT) and NTT highlights a growing ecosystem of partnerships designed to provide the testing grounds (labs and POVs) necessary for enterprises to validate these autonomous workflows. The ultimate consequence is a significant reduction in noise—where thousands of alerts are distilled into a handful of actionable, or even self-resolving, insights.

Conclusion

Agentic AIOps marks the beginning of the autonomous enterprise. By combining the deep visibility of infrastructure monitoring with the sophisticated reasoning of generative AI, organizations can finally address the scale and speed requirements of modern digital services. While the technology is revolutionary, its success remains rooted in the fundamentals: high-quality data, clear governance, and a phased approach to building autonomous trust.

Links:

PostHeaderIcon [AWSReInvent2025] Supercharging DevOps with AI-Driven Observability: The Next Frontier in SRE

Lecturer

Elizabeth Fuentes is a Senior Developer Advocate at Amazon Web Services (AWS), specializing in the intersection of Artificial Intelligence and DevOps practices. With extensive experience in cloud architecture and software engineering, Elizabeth focuses on how Generative AI can streamline complex CI/CD pipelines and enhance Site Reliability Engineering (SRE). She is a key contributor to AWS educational initiatives, having co-developed advanced courses on AI-driven automation. Joining her is Laas Alina, a software architect and open-source enthusiast who focuses on implementing multi-agent systems and the Model Context Protocol (MCP) to solve observability challenges at scale.

Abstract

As software systems grow increasingly distributed and complex, traditional observability—centered on manual log analysis and reactive dashboards—is becoming insufficient. This article explores the paradigm shift toward AI-driven observability, where Generative AI serves not just as a query tool, but as an active participant in failure detection, correlation, and resolution. By leveraging Amazon Bedrock and Amazon Q, organizations can transition from “reactive” to “predictive” DevOps. The discussion analyzes the methodology of building AI agents that simulate architectural stress, automatically explain multi-layered failures, and provide traceable, actionable recommendations. We examine the implementation of the Model Context Protocol (MCP) in establishing sophisticated multi-agent systems (MAS) that transform raw data into contextual understanding, ultimately reducing the Mean Time to Resolution (MTTR) and enhancing systemic resilience.

The Evolution of Observability: From Metrics to Contextual Understanding

The traditional pillars of observability—metrics, logs, and traces—provide the “what” of a system’s state but often fail to provide the “why” in real-time. In high-velocity DevOps environments, the sheer volume of telemetry data can overwhelm human operators, leading to “alert fatigue” and delayed responses to critical incidents. Elizabeth posits that the integration of Generative AI marks the fourth pillar of observability: Contextual Intelligence. This evolution moves the industry beyond simple threshold-based monitoring toward systems that understand the semantic relationship between a failed deployment, a spike in latency, and a specific line of code.

By utilizing Large Language Models (LLMs) through Amazon Bedrock, DevOps teams can ingest vast amounts of unstructured log data and receive summaries that highlight anomalies that might be missed by traditional regex-based filters. The methodology involves training the AI to recognize “normal” operational patterns and identifying deviations not just by value, but by the intent of the system’s behavior. This contextual layer allows for a more nuanced interpretation of system health, where the AI can distinguish between a benign resource spike and a precursor to a cascading failure.

Architecting AI Agents for Predictive Troubleshooting

The transition to AI-driven observability is characterized by the deployment of “Micro-agents”—specialized AI entities designed to handle specific segments of the DevOps lifecycle. These agents operate within a Multi-Agent System (MAS), where they collaborate to solve complex incidents. For instance, a “Monitoring Agent” might detect a performance degradation and immediately trigger a “Diagnosis Agent” to correlate the event with recent CI/CD pipeline changes.

Elizabeth and Laas Alina emphasize the importance of the Model Context Protocol (MCP) in this architecture. MCP acts as the communication backbone, allowing agents to share context without losing the “lineage” of a decision. When an AI agent recommends a specific architectural change or a rollback, it must provide clear traceability. This is crucial for maintaining trust in automated systems. The agents do not operate in a vacuum; they interact with tools like Amazon Q to provide developers with instant explanations of failures directly within their Integrated Development Environment (IDE) or chat interface.

// Example of an AI-driven Observability Agent Configuration
agent:
  name: "IncidentDiagnosticAgent"
  provider: "AmazonBedrock"
  model: "claude-3-sonnet"
  capabilities:
    - log_analysis
    - metric_correlation
    - trace_summarization
  mcp_config:
    protocol_version: "1.0"
    shared_context: "deployment_metadata"
  safety_guardrails:
    - max_token_usage: 4000
    - human_in_the_loop_required: true

Transforming CI/CD through Generative AI and Simulation

Beyond reactive troubleshooting, AI-driven observability empowers proactive system design. One of the most innovative concepts discussed is the use of AI agents to simulate “stress-test” scenarios within a digital twin of the production environment. These agents can intentionally inject failures—similar to Chaos Engineering—and then observe how the observability stack responds. This creates a feedback loop where the AI helps engineers identify “blind spots” in their monitoring before a real incident occurs.

Furthermore, Generative AI transforms the CI/CD pipeline by automatically generating “failure explanations.” Instead of a developer sifting through a 5,000-line build log, Amazon Q can provide a concise summary: “The build failed because the new database schema in commit X is incompatible with the connection pool settings in environment Y.” This level of automated insight accelerates the “inner loop” of development, allowing engineers to focus on innovation rather than infrastructure archeology.

The Human-AI Partnership: Strategic Implications

A common concern in the industry is the replacement of human engineers by AI. However, Elizabeth argues that the future belongs to the “augmented engineer.” AI is a force multiplier that automates the repetitive, “drudge work” of observability—log parsing and initial triage—allowing human experts to focus on high-level strategy and complex architectural decisions. The goal is to transform teams from being “reactive” (fighting fires) to “proactive” (preventing fires).

Implementing these systems requires a cultural shift toward AI-literacy within DevOps teams. Organizations must establish safety guardrails to ensure that AI-driven recommendations are validated and that automated actions (like auto-remediation) have clear rollback paths. By embracing AI as a strategic tool, DevOps and SRE teams can achieve a level of operational excellence that was previously unattainable, ensuring that as systems grow in scale, their reliability grows in parallel.

Links:

PostHeaderIcon [AWSReInvent2025] Accelerating Enterprise Modernization: The Architecture of Composable AI Agents

Lecturer

Mortaza Chowri is the Head of Product Management for the AWS Transform team, where he leads the development of next-generation tools for complex workload migration. He is an expert in leveraging generative AI to automate technical debt reduction for large-scale enterprises. Joining him are Alexi and Ravi, who serve as senior architects within the AWS Transform division, specializing in agentic AI implementation and the creation of composable system frameworks. The session also features strategic insights from the leadership team at Capgemini, who collaborate with AWS to deliver industry-specific modernization solutions for global banking and automotive clients.

Abstract

Enterprise modernization is frequently paralyzed by the extreme complexity of legacy systems, particularly decades-old mainframes and aging Windows-bound .NET applications. This article explores the innovative framework of AWS Transform, a centralized service that utilizes “Agentic AI” to automate and streamline the migration process. The methodology centers on the concept of composability, which allows AWS partners to integrate their proprietary industry knowledge and specialized tools with foundational AI agents. By utilizing a sophisticated chat-based interface and automated business rule extraction, the platform enables a seamless transition from legacy COBOL and .NET Framework 4.x to modern, cloud-native architectures. The analysis demonstrates how these composable agents create a continuous feedback loop that significantly reduces manual effort, improves documentation, and ensures business logic remains intact during high-risk migrations.

Context: The Burden of Technical Debt and Knowledge Atrophy

Many of the world’s most critical systems, particularly in finance and manufacturing, are still dependent on infrastructure built in the late 20th century. These legacy environments present three primary obstacles that prevent organizations from achieving modern agility. First, knowledge atrophy has become a critical risk, as the original architects of these mainframe systems have often retired, leaving behind “black box” applications that lack contemporary documentation. Second, the technical debt associated with older languages like COBOL is immense, as these systems were never designed to leverage modern cloud features such as serverless compute or elastic auto-scaling.

Third, the mission-critical nature of these systems creates a state of risk aversion, where the fear of breaking a core business process during a manual rewrite often leads to stagnation. AWS Transform was specifically developed to break this cycle of inertia. By providing a unified experience that integrates discovery, assessment, and modernization into a single platform, AWS allows enterprises to view their legacy code as an asset to be reimagined rather than a liability to be feared.

Methodology: Agentic AI and the Composable Framework

The core technical innovation of AWS Transform is the transition from static point solutions to a dynamic, “unified experience” powered by specialized AI agents. These agents are designed to perform complex technical tasks with a level of autonomy that far exceeds traditional automation scripts. The methodology is built upon several key pillars of agentic behavior. Discovery agents are tasked with automatically mapping technical artifacts, such as physical servers and complex database schemas, to their optimal cloud-native equivalents.

Modernization agents, specifically those tuned for mainframe environments, perform the difficult work of extracting business rules from legacy code. This process generates comprehensive documentation that allows current engineers to “comprehend” the underlying logic of systems they did not build. The most transformative aspect of this methodology is its composability for partners. AWS provides the foundational intelligence and large language models, while partners such as Capgemini can “compose” these with their own specialized knowledge bases and custom transformation rules. This enables the creation of industry-specific agents, such as a modernization assistant specifically optimized for banking regulations or complex automotive production logic.

Technical Analysis of Mainframe Rule Extraction

The implementation of these agents in real-world scenarios, particularly through the collaboration with Capgemini, highlights a sophisticated “forward engineering” approach. In this workflow, the AI agents first scan the legacy code to identify core business logic and immutable rules. This extraction phase is critical because it ensures that while the code is updated, the essential business functions remain perfectly intact. Following extraction, the reimagination phase begins, where these rules are integrated into a modern architecture that meets cloud-native standards for security and performance.

Practitioners interact with these systems through a chat experience within the AWS Transform interface, allowing them to query both the AI agents and integrated domain experts directly. This interaction model democratizes the modernization process, making it accessible to developers who may not have expertise in COBOL but are proficient in modern languages like Java or Python. The platform serves as a bridge, translating the “what” of legacy business logic into the “how” of modern cloud execution.

Outcomes: Efficiency, Consistency, and Continuous Learning

The deployment of composable AI agents has fundamentally altered the economics and speed of enterprise modernization. By automating the most labor-intensive parts of code comprehension and translation, organizations have reported a reduction in manual effort by as much as 80%. This allows teams to focus on high-value innovation rather than the repetitive task of line-by-line code migration. Furthermore, the platform ensures architectural consistency across a large organization, preventing the fragmentation that often occurs when different teams use varying migration tools.

One of the most significant consequences of this approach is the continuous improvement of the agents themselves. Every modernization task performed through the platform provides feedback data that enhances the underlying AI models. As these agents encounter more diverse enterprise environments, their ability to handle edge cases and complex business rules grows exponentially. This creates a virtuous cycle where each successful migration makes the next one faster and more reliable, effectively solving the problem of knowledge atrophy for the long term.

Conclusion

The shift toward agentic AI and composable architectures represents a milestone in the evolution of enterprise IT. AWS Transform provides a robust framework that allows organizations to tackle their most daunting legacy challenges with a level of confidence and speed that was previously impossible. By allowing partners to integrate their unique industry expertise into a centralized AI system, AWS has created a scalable ecosystem that transforms modernization from a risky, multi-year endeavor into a manageable and continuous strategic process.

Links: