Recent Posts
Archives

Posts Tagged ‘AI’

PostHeaderIcon [GoogleIO2026] Google I/O 2026 Developer Keynote: Deep Dive into Agentic Workflows, Infrastructure, and Cross-Platform Systems

Lecturer

Josh Woodward, Logan Kilpatrick, Paige Bailey, Anshul Bhagi, Kevin Moore, Florina Muntenescu, Adarsh Fernando, Yuna Kravets, and Matthias Bynens presented the latest ecosystem updates across Google AI Studio, Google Antigravity, Android, and Chrome.

Abstract

This article provides a comprehensive technical analysis of the systems, runtime harnesses, developer tools, and platform APIs unveiled during the Google I/O 2026 Developer Keynote. Key updates include the launch of Gemma 4, managed agents in the Gemini API with remote sandboxing, Google Antigravity 2.0 (featuring dynamic subagents, cron scheduled tasks, and CLI integration), native agentic workflows in Android Studio and the Android CLI, and the evolution of the Agentic Web via Web MCP, Modern Web Guidance, and Chrome DevTools for agents.

Managed Agents Runtime and AI Studio Ecosystem

The transition toward goal-driven autonomous systems requires orchestration layers that abstract compute isolation and tool access. Google expanded its developer runtime capabilities through open-source foundation models and managed execution infrastructure.

Open Model Advances: Gemma 4

Gemma 4 was released under an Apache 2 license, designed specifically for advanced reasoning, local intelligence, and on-device agentic execution. Key achievements include:

  • Deployment Versatility: Compact footprint capable of running offline on mobile devices, robotics systems, and satellite hardware.
  • Ecosystem Adoption: Surpassed 100 million downloads in its first month, propelling total cumulative Gemma series downloads past 500 million.
+-----------------------------------+
|      Gemma Series Download Metric |
+-----------------------------------+
| Initial Month (Gemma 4):  100M    |
| Cumulative Gemma Series: >500M    |
+-----------------------------------+

Managed Agents in Gemini API & Interactions API

Building on the Interactions API introduced in late 2025, Google introduced managed agents directly within the Gemini API.

+---------------+      API Call     +------------------+
| User Request  | ----------------> | Gemini Managed   |
+---------------+                   | Agent Runtime    |
                                    +--------+---------+
                                             |
                                    Provisions & Isolates
                                             |
                                             v
                                    +------------------+
                                    | Remote Linux Sandbox|
                                    | (Compute Environment)|
                                    +------------------+

  • Remote Linux Sandboxing: Every managed agent call provisions a secure, isolated remote Linux execution environment in Google Cloud. The platform handles state provisioning, runtime dependencies, and compute isolation.
  • Declarative Markdown Configuration: Skills, custom instructions, tools, and memory parameters are defined using standard .md files (e.g., agents.md), allowing declarative agent engineering without custom orchestration logic.“`
+-----------------------------------+
|   Managed Agent Modular Architecture   |
+-----------------------------------+
| Skill Configuration (Markdown)   |
|  - Research (Web Fetching/APIs)   |
|  - Scriptwriting / Text Gen      |
|  - Multi-Voice TTS Synthesis     |
|  - Lyria Music Generation        |
|  - Audio Mixing & Master Output  |
|  - Nano Banana Asset Generation  |
+-----------------------------------+

AI Studio Workflow & Deployment Enhancements

Google AI Studio updated its visual platform to support rapid prototyping and multi-platform deployment:

  • One-Click Cloud Run Deployment: Instant deployment of web applications to live Cloud Run URLs with zero credit card setup for new developers.
  • Full-Stack Integrations: Native bindings for Firebase, Firestore, Google Workspace (Docs, Gmail, Calendar), and Google Search.
  • Native Android App Generation: Direct synthesis of Kotlin codebase previews within an embedded Android emulator inside AI Studio. Includes direct APK delivery to physical USB-tethered devices and automated deployment pipelines to Google Play Store test tracks.
  • AI Studio Mobile App: Pre-registration launched for a dedicated iOS/Android application bringing prompt-to-app workflows to mobile form factors.
  • Antigravity Portability: One-click full filesystem export from Google AI Studio into local Antigravity environments without state loss.

Google Antigravity 2.0 and Agent Orchestration

Google Antigravity 2.0 shifts developer interactions from command line completion to asynchronous, multi-agent execution environments.

                    +-----------------------+
                    | Anti-Gravity 2.0      |
                    | Mission Control       |
                    +-----------+-----------+
                                |
     +--------------------------+--------------------------+
     |                          |                          |
+----+-----+               +----+-----+               +----+-----+
| Subagent |               | Subagent |               | Subagent |
| (Task A) |               | (Task B) |               | (Task C) |
+----+-----+               +----+-----+               +----+-----+
     |                          |                          |
Worktree 1                 Worktree 2                 Worktree 3

Core Architecture and Features

  • Multi-Worktree Concurrency: Run simultaneous agents in separate Git worktrees across disparate projects without file collisions.
  • Dynamic Subagents: Autonomous creation of specialized worker subagents (e.g., QA, data science, refactoring) executing in parallel.
  • Scheduled Tasks (Cron Autopilot): Native support for standard cron syntax allowing proactive background agent execution (e.g., automated morning PR summarization or hourly cloud infrastructure health checks).
  • Antigravity SDK & Enterprise Cloud Binding: Programmatic developer control over agent harnesses and enterprise project binding under standardized enterprise security terms.
  • Domain Skills Bundles: Pre-packaged capabilities for specialized domains, starting with the Scientific Skill Bundle for accelerating biology, health, and research tasks.

Command Line Integration: Antigravity CLI

The unified Antigravity CLI merges the legacy Gemini CLI into the standalone Antigravity runtime:

  • Provides an identical agent harness and model access within terminal environments, supporting custom themes, keybindings, and headless SSH sessions.
  • Features interactive side-channel commands like /btw to fork quick model queries without corrupting the main conversation or context window.
+-----------------------------------+
|     Gemma 4 Fine-Tuning Bench      |
+-----------------------------------+
| Dataset: Prompt -> Bash Mapping   |
| Technique: LoRA Parameter Efficient|
| Environment: Remote GPU VM via CLI|
| Deployment: Local Ollama/SGLang   |
+-----------------------------------+

Android Platform Architecture & Studio Integrations

Native Android development receives native agent capabilities via the Android CLI and Android Studio tooling integration.

+-----------------------------------+
|     Android CLI Agent Architecture|
+-----------------------------------+
| Knowledge Base + Open Source Skills|
|                |                  |
|                v                  |
| Context-Aware Token Reduction     |
| (70% Token Cut / 3x Exec Speed)   |
|                |                  |
|                v                  |
| Android Studio IDE Hook Integration|
+-----------------------------------+

Android CLI & Knowledge Base

The built-in Android CLI exposes SDK management, project instantiation, UI compilation, and device deployment directly to autonomous agents.

  • Android Knowledge Base & Open-Source Skills: Provides models with up-to-date best practices (e.g., XML to Jetpack Compose migrations, Jetpack Navigation 3, edge-to-edge layouts).
  • Token Efficiency: Benchmarks demonstrate a 70% reduction in context token consumption and a 3x speedup in task completion times when using guided Android skills.
+-----------------------------------+
| Jetpack Compose Glimmer XR Engine |
+-----------------------------------+
| Hybrid Execution Architecture      |
|  - On-Device: Gemini Nano 4       |
|  - Cloud Fallback: Firebase AI    |
+-----------------------------------+

IDE Optimizations and Quality Tooling

  • R8 Configuration Analyzer Skill: Automated audit of ProGuard/R8 keep rules and build scripts to enable full-mode shrinking, reduce app size, and eliminate Application Not Responding (ANR) occurrences.
  • App Links Assistant Integration: Automated parsing of web URLs to generate activity mapping logic, deep-linking intent filters, and unit test validations.
  • Android Device Streaming Expansion: Support for real hardware target streaming, including the Samsung Galaxy S26 Ultra.
+-----------------------------------+
| Native Cross-Platform Migration   |
+-----------------------------------+
| Source: iOS / Web / React Native  |
| Engine: Android Studio Assistant  |
| Pipeline: Storyboard -> Jetpack UI|
| Target: Kotlin Multiplatform (KMP)|
+-----------------------------------+

Agentic Web, Chrome DevTools, and Modern Web Standards

The web platform is undergoing a fundamental transformation to ensure sites are fully readable, actionable, and testable by browser agents.

+-----------------------------------+
| Modern Web Baseline Standards     |
+-----------------------------------+
| Mapping Target: 100% Cross-Browser|
| Modern Web Guidance: Token Efficient|
| Benchmark Gain: +37% Pass Rate    |
+-----------------------------------+

Web Model Context Protocol (Web MCP)

Web MCP is an experimental browser standard proposed to expose site capabilities directly to client-side LLM agents.

+-----------------+                      +-------------------+
| Web Page / App  |  Registers Schemas   | Gemini in Chrome  |
| (React/Angular) | -------------------> | (Browser Agent)   |
+--------+--------+                      +---------+---------+
         |                                         |
         |        Executes JavaScript Tool Calls   |
         + <---------------------------------------+

  • Imperative Web Tools: Developers expose programmatic JavaScript tools and schema parameters (e.g., updateCarConfiguration) directly to the browser runtime.
  • Origin Trial Target: Experimental Web MCP APIs launch in Chrome 149, with native execution support in Chrome’s side-panel agent.

Chrome DevTools for Agents

To close the execution-feedback loop for coding agents, Chrome introduced DevTools integration optimized for autonomous systems:

  • Agentic Browsing Audits in Lighthouse: Evaluates Web MCP tool registrations, llms.txt discovery manifests, declarative form labels, and accessibility tree ARIA roles.
  • Autonomous Feedback Loop: Agents connect directly via the Model Context Protocol (MCP), execute runtime audits, analyze error stacks, patch source code, and verify fixes autonomously without developer copy-pasting.
+-----------------------+     Runs Audit     +-----------------------+
| Chrome DevTools Agent | -----------------> | Lighthouse Engine     |
+-----------^-----------+                    +-----------+-----------+
            |                                            |
            |            Emits Error/ARIA Log            |
            +<-------------------------------------------+
            |
    Applies Source Fix
            |
            v
+-----------------------+
| Local Project Code    |
+-----------------------+

Hardware-Accelerated Web Graphics: HTML in Canvas

The HTML Canvas API now supports direct rendering of live, interactive DOM elements inside Canvas contexts (including 3D WebGL scenes).

  • Accessibility and Interactivity: Rendered DOM elements remain fully selectable, searchable, accessible to assistive technologies, translatable, and compatible with browser autofill features.

Ecosystem Initiatives and Pricing

Google introduced several developer support mechanisms and enterprise tiers to scale agentic deployment:

  • Build with Gemini X Prize Hackathon: A global developer competition featuring $2,000,000 in total prizes for real-world impact projects leveraging Gemini APIs.
  • Google AI Ultra Plan: A $100 per month developer tier providing elevated rate limits, enterprise platform features, and $100 in bonus Antigravity runtime credits.

Links:

PostHeaderIcon [MiamiJUG] Specialization and Efficiency: The Future of Distilled Models and MoE

Lecturer

Frank Greco is a distinguished Java Champion and enterprise architect with a deep focus on AI, Cloud, and Edge computing. As a senior consultant and long-standing educator, he chairs the NYJavaSIG and has co-authored industry standards such as JSR #381. Frank is dedicated to helping developers navigate the practical implementation of machine learning within enterprise ecosystems.

Abstract

As generative AI moves from experimental prototypes to enterprise production, the focus has shifted from monolithic models to specialized architectures. This article analyzes two critical trends: Distilled Models and Mixture of Experts (MoE). By exploring how large models can “teach” smaller, more efficient versions and how sub-networks can be orchestrated to handle niche tasks, this study provides a roadmap for building cost-effective, high-performance AI applications in memory-constrained environments.

The Methodology of Model Distillation

The current evolution of AI prioritizes efficiency and latency over raw parameter count. Model distillation is a process where a large, high-parameter model (the “Teacher”) is used to train a significantly smaller model (the “Student”).

The technical process involves:

  1. Reasoning Extraction: The teacher model is prompted to solve problems using Chain of Thought (CoT) reasoning.
  2. Pattern Learning: The student model is trained on the teacher’s thought process and step-by-step logic.
  3. Optimization: The resulting student model—such as the DeepSeek variants—retains much of the reasoning capability of the larger model while requiring significantly less memory and providing faster response times.

This is particularly relevant for Java developers who need to deploy AI features in environments where the infrastructure costs of running a massive LLM would be prohibitive.

Mixture of Experts (MoE) Architecture

Beyond distillation, the industry is transitioning toward “Mixture of Experts” (MoE) architectures. Instead of one massive, uniform neural network, an MoE system consists of a collection of specialized sub-networks.

In this configuration, a “router” analyzes the incoming prompt and determines which “expert” sub-network is best suited to answer. For instance, a technical query about Java garbage collection would be routed to a code-specialized network, whereas a question about financial regulation would go to a legal-specialized expert. This approach ensures higher precision and reduces the total active parameters needed for a single query, leading to more efficient processing at scale.

Conclusion: The Developer as Orchestrator

The emergence of these specialized architectures changes the role of the enterprise developer. Rather than simply querying a single general-purpose model, developers must now act as orchestrators, selecting the right combination of distilled models and expert networks for their specific domain. By understanding these architectural shifts, engineers can build AI-integrated systems that are both powerful and economically viable for large-scale production.

Links:

PostHeaderIcon [GoogleIO2026] Google I/O 2026 Keynote: Advances in Multimodal AI, Agentic Workflows, and Spatial Computing

Lecturer

Sundar Pichai is the Chief Executive Officer of Alphabet Inc. and its subsidiary Google. Holding degrees from the Indian Institute of Technology Kharagpur, Stanford University, and the Wharton School of the University of Pennsylvania, he has overseen the organization’s strategic shift toward an AI-first approach over the past decade.

Abstract

This article analyzes the technological breakthroughs, system architectures, and product paradigms presented at the Google I/O 2026 Keynote. Key announcements include the introduction of the Gemini 3.5 model family, the Gemini Omni multimodal world model, the Google Antigravity 2.0 agent-first development platform, and the integration of autonomous agents across Search, Workspace, and Android XR hardware. The technical, economic, and security implications of these innovations are examined in detail.

Infrastructure Scale and Custom Silicon Evolution

Scaling state-of-the-art artificial intelligence models requires unprecedented investments in compute infrastructure and specialized hardware architectures. Capital expenditure has escalated significantly, transitioning from 31 billion dollars annually in 2022 to an estimated range of 180 to 190 billion dollars. This dramatic funding increase underscores the foundational compute demands required to serve thousands of trillions of tokens across billions of global consumer and enterprise touchpoints.

A central driver of this infrastructure strategy is the eighth generation of custom Tensor Processing Units (TPUs). Google introduced a dual-chip paradigm tailored for distinct machine learning workloads:

  • TPU 😯 (Training Optimized): Engineered specifically for large-scale pre-training, delivering nearly three times the raw computing power of previous iterations.
  • TPU 8i (Inference Optimized): Architected to minimize latency and improve energy efficiency, delivering up to two times better performance per watt.
+-----------------------------------+
|      Google TPU Generation 8      |
+-----------------+-----------------+
| TPU 8O          | TPU 8i          |
| (Training)      | (Inference)     |
+-----------------+-----------------+
| * 3x Power      | * Low Latency   |
| * Distributed   | * ~1500 Tok/s   |
| * Multi-site    | * 2x Perf/Watt  |
+-----------------+-----------------+

To bypass the physical limits of individual data center facilities, the Jackson Pathways framework allows distributed pre-training across multiple global sites simultaneously. In inference benchmarks, next-generation Flash models executing on TPU 8i silicon achieved output processing rates approaching 1,500 tokens per second. Overall platform usage expanded to 3.2 quadrillion tokens per month, driven by over 8.5 million active developers.

+-----------------------------------+
|      Monthly Token Trajectory     |
+-----------------------------------+
| 2024: 9.7 Trillion Tokens         |
| 2025: 480 Trillion Tokens         |
| 2026: 3.2 Quadrillion Tokens      |
+-----------------------------------+

Frontier Multimodal Models and World Simulation

The frontier of generative modeling is shifting from static media generation to dynamic world simulation. The flagship Gemini Omni model unifies core large language model reasoning with specialized generative media models such as Veo, Nano Banana, and Genie.

       +--------------------+
       | Gemini Core Engine |
       +---------+----------+
                 |
     +-----------+-----------+
     |           |           |
+----+-----+ +---+------+ +--+-----+
|   Veo    | |   Nano   | | Genie  |
| (Video)  | |  Banana  | | (Sims) |
+----+-----+ +---+------+ +--+-----+
     |           |           |
     +-----------+-----------+
                 |
       +---------v----------+
       |    Gemini Omni     |
       |   (World Model)    |
       +--------------------+

Gemini Omni functions as a world model capable of understanding kinetic energy, gravitational mechanics, three-dimensional geometry, and physical interactions. It processes heterogeneous inputs—text, raster images, structured data, and video streams—to generate high-fidelity, interactive outputs.

To address the proliferation of synthetic media, Google expanded its digital provenance framework. The SynthID watermarking technology—which has marked over 100 billion images and videos alongside 60,000 years of audio assets—is complemented by explicit Content Credentials. Integrated into Google Search and Chrome via Circle to Search and context menu controls, these mechanisms verify whether content originated from physical hardware sensors or underwent generative editing.

Agentic Development Frameworks and Autonomous Systems

Agentic capabilities represent a fundamental shift from assisted output creation to goal-driven autonomous execution. Gemini 3.5 Flash serves as the foundational model for high-speed agentic tasks, demonstrating superior latency-to-intelligence ratios and performing four times faster than previous frontier models.

Google Antigravity 2.0

The agent-first software development platform, Antigravity 2.0, reorganizes developer workflows around multi-agent orchestration, asynchronous execution, and subagent teamwork. Key system primitives include:

  • Subagent Networks: Division of complex engineering goals into parallel subtasks.
  • Execution Hooks and Harnesses: Sandboxed environments providing file read/write, terminal command invocation, and automated unit test verification.
  • CLI and Native SDK Integrations: Programmatic control binding into local development environments, Android, Firebase, and Google AI Studio.

In stress-testing evaluations, an autonomous network of 93 Antigravity subagents executed over 15,000 model requests and processed 2.6 billion tokens over a 12-hour period to construct a fully functional operating system kernel—including memory management, task scheduling, and file systems—from scratch.

+-----------------------------------+
|  Antigravity Autonomous OS Build  |
+-----------------------------------+
| Subagents Active:  93             |
| Model Requests:   >15,000         |
| Tokens Processed:  2.6 Billion    |
| Build Duration:    12 Hours       |
| Total API Cost:   <$1,000         |
+-----------------------------------+

Consumer Agent Integration: Gemini Spark

For end-user workflows, Gemini Spark introduces persistent background execution environments running on dedicated virtual machines in Google Cloud. Utilizing the Model Context Protocol (MCP) and the Antigravity agent harness, Spark handles multi-step, asynchronous directives without requiring active user sessions.

Agent commerce protocols extend these execution capabilities to financial transactions:

  • Universal Commerce Protocol (UCP): An open-source communication layer standardizing product search, inventory mapping, and checkout across diverse merchant platforms.
  • Agent Payments Protocol (AP2): Security protocols utilizing cryptographic digital mandates and strict spending boundaries to execute authenticated transactions on behalf of users.
+---------------+
| User Intent   |
+-------+-------+
        |
        v Cryptographic Mandate
+---------------+
| Agent (AP2)   |
+-------+-------+
        |
        v Validated Boundary
+---------------+
| Google Pay    |
+-------+-------+
        |
        v Digital Trail
+---------------+
| Merchant      |
+---------------+

Agentic Search, Generative Interfaces, and Spatial Computing

Google Search has transitioned into a native AI Search engine, consolidating traditional indexing with real-time generative capabilities.

Dynamic Generative UI

Leveraging Gemini 3.5 Flash within containerized execution sandboxes, Search dynamically designs and renders interactive user interfaces on the fly. When handling complex conceptual queries, the system writes layout code, computes parameters, and renders custom widgets or stateful micro-applications directly within the search results stream.

User Query
    |
    v
Intent Analysis
    |
    v
Agent Harness (Antigravity)
    |
    v
Generates UI & Code
    |
    v
Dynamic Rendered Visual

Spatial Computing and Intelligent Eyewear

In spatial computing, Android XR expands beyond headsets to intelligent eyewear. Audio glasses featuring integrated Gemini models deliver context-aware, heads-up interactions via directional audio drivers. Operating in tandem with personal intelligence APIs, these wearables interpret real-time environmental context, facilitate hands-free navigation, execute app workflows via voice, and interface with smartwatches for compact visual previews.

Scientific Discovery Engine and Singularitarian Horizons

The application of artificial intelligence to physical sciences represents a pivotal paradigm shift. Gemini for Science consolidates predictive tools, code synthesis, paper digestion, and hypothesis formulation into unified laboratory workflows.

Central to this scientific strategy is high-performance dynamic simulation. Alpha Earth Foundations models planetary mechanics as a digital twin to predict climate anomalies, deforestation, and agricultural vulnerability. In atmospheric science, Weather Next superseded classical numerical fluid dynamics, accurately forecasting Category 5 hurricane trajectories days prior to landfall.

+-----------------------------------+
|  Alpha Earth & Weather Next Engine|
+-----------------------------------+
| Physical Data Assimilation        |
|                |                  |
|                v                  |
| AI Twin Simulation Layer          |
|                |                  |
|                v                  |
| Predictive Early Alerts           |
+-----------------------------------+

In molecular biology, Isomorphic Labs leverages deep generative architectures to model molecular interactions at atomic precision. Moving beyond static target predictions toward preclinical drug discovery, the platform actively accelerates therapeutic candidate synthesis for oncology and autoimmune pathologies. These systems signify a systematic transition toward digital-speed empirical research.

Links:

PostHeaderIcon [VoxxedDaysBucharest2026] Optimizing LLM Inference on Kubernetes: Abdel Sghiouar on Practical Techniques for the Rest of Us

Lecturer

Abdel Sghiouar is a Developer Advocate at Google Cloud with deep expertise in cloud-native technologies, Kubernetes orchestration, and AI/ML workload optimization. Drawing from a robust background in infrastructure engineering and open source contributions, Abdel helps organizations design, deploy, and tune complex AI applications for production environments across diverse infrastructures.

Abstract

While major cloud providers and hyperscalers leverage virtually unlimited computational resources, the majority of organizations face significant constraints when operationalizing Large Language Models. Abdel Sghiouar presents a comprehensive set of practical strategies for optimizing LLM inference workloads on Kubernetes. The session systematically addresses container and model optimization techniques, accelerator management, data persistence and storage considerations, networking and intelligent load balancing, and advanced observability practices. Emphasis is placed on open-source tools and architectural patterns that deliver meaningful cost-performance improvements adaptable to on-premises, hybrid, and public cloud deployments.

Understanding LLM Inference Characteristics and Challenges

Large Language Models continue their rapid evolution in both scale and sophistication. Architectural innovations such as mixture-of-experts (MoE) enable dynamic activation of specialized sub-networks, while multi-modal capabilities process diverse inputs including text, images, audio, and video. Expanded context windows support richer interactions but demand substantial memory resources.

Inference execution comprises two primary phases with contrasting characteristics: the prefill stage (encoding input tokens, predominantly compute-bound) and the decode stage (token generation, typically memory-bound). KV (key-value) caching optimizes conversational flows by preserving intermediate states, avoiding redundant prefill computations for subsequent messages.

Deployment topologies vary considerably. Single-host single-accelerator setups predominate for local development and experimentation (e.g., using Ollama). Single-host multi-accelerator configurations require model sharding across GPUs within one machine. Multi-host distributed deployments introduce complex requirements for high-bandwidth, low-latency interconnects to maintain coherent context across nodes. Each topology presents distinct challenges regarding scalability, fault tolerance, and operational complexity.

Container, Model, and Storage Optimizations

Inference serving runtimes and model artifacts generate exceptionally large container images, frequently exceeding several gigabytes prior to incorporating weights. Conventional optimization strategies like multi-stage builds or native compilation (e.g., GraalVM) prove inadequate for these workloads.

Distributed caching solutions such as Spiegel provide cluster-wide image and model artifact caching, substantially reducing repeated pulls from external registries. Kubernetes-native features enabling containers as volumes allow separate packaging of models, which can then be mounted efficiently onto serving runtimes. When combined with caching layers, these approaches dramatically accelerate cold starts.

Quantization techniques offer another lever, reducing numerical precision (e.g., FP16 to INT8 or lower) to decrease memory footprints while preserving sufficient accuracy for many applications. Careful selection of quantization levels based on task sensitivity balances performance and quality.

Accelerator Management and Dynamic Resource Allocation

Kubernetes has supported GPU scheduling through device plugins for several years. However, static device configurations struggle with real-world constraints including accelerator scarcity and heterogeneous hardware fleets.

Dynamic Resource Allocation, matured in recent Kubernetes versions, introduces flexible resource claiming based on abstract characteristics rather than rigid device specifications (e.g., requesting “NVIDIA GPU with minimum 30GB memory and specific core count”). This enables more efficient scheduling across mixed clusters and better utilization rates.

Integration with cluster autoscalers allows on-demand provisioning, addressing both availability gaps and cost optimization by scaling resources precisely to workload demands. Platform operators describe device inventories; application teams specify requirements, with the scheduler performing intelligent matching.

Networking, Load Balancing, and Observability Considerations

LLM traffic profiles differ markedly from conventional web workloads. Requests exhibit high variability in size and computational intensity (simple text queries versus multi-modal inputs), while responses frequently involve streaming token generation. Standard round-robin load balancing produces inefficient distributions, with certain backends becoming overloaded while others remain underutilized.

The Kubernetes Gateway API, augmented with custom endpoint selection logic, supports sophisticated routing decisions based on request attributes extracted from bodies (model identifier, input modality, streaming requirements) combined with real-time backend telemetry. This facilitates intelligent traffic steering, prioritization of business-critical workloads, and maintenance of sticky sessions necessary for coherent streaming interactions.

Comprehensive observability must encompass prefill and decode phase latencies, KV cache hit rates, token generation throughput, GPU utilization, and end-to-end request metrics. Integration with Prometheus, Grafana, and specialized LLM monitoring solutions provides actionable insights for capacity planning and bottleneck identification.

Practical Patterns and the LLM-D Project

The LLM-D initiative, hosted under the Linux Foundation with contributions from Google, IBM, NVIDIA, and additional partners, aggregates architectural patterns, performance benchmarks, and reference implementations for production-grade inference. Key elements include optimized prefill/decode separation, advanced routing logic often leveraging engines like vLLM, and comprehensive guidance for multi-node deployments.

A holistic, layered optimization strategy proves most effective: infrastructure-level improvements (caching, persistent volumes), platform capabilities (dynamic scheduling, intelligent networking), and application-level choices (model quantization, serving engine selection). Organizations without hyperscale resources can still achieve competitive efficiency and scalability through disciplined application of these patterns.

Links:

PostHeaderIcon [AWSReInvent2025] The Next Frontier in Financial Systems: Architecting Transformer-based Foundation Models for Real-Time Payments

Lecturer

Sudeep Kalindi is a Principal Solution Architect at Amazon Web Services (AWS), where he focuses on building scalable AI and machine learning solutions for the global financial services industry. With a deep expertise in high-frequency transaction systems and cloud infrastructure, Sudeep advises major financial institutions on modernizing their fraud detection and personalization engines using advanced neural network architectures.

Pahal Patangia is the Global Head of Business for the Payments Industry at NVIDIA. He has spent nearly five years at NVIDIA accelerating the adoption of AI and accelerated computing within the payments ecosystem. Pahal works closely with banks, fintechs, and payment processors to deploy large-scale foundation models that transform transactional data into real-time business value.

Abstract

As digital transactions explode in volume and complexity, traditional rule-based and machine learning models are reaching their limits in combating sophisticated fraud and providing personalized customer experiences. This article examines the emergence of transformer-based foundation models as the “next frontier” for financial systems. Unlike prior models that treated transactions as isolated events, transformers excel at capturing long-term dependencies and sequential patterns in tabular transactional data. The discussion details the technical advantages of “attention” mechanisms in finance, the role of NVIDIA’s accelerated computing in training these massive models, and the deployment strategies on AWS that enable real-time inference. By integrating tabular foundation models with Graph Neural Networks (GNNs), financial institutions can achieve unprecedented accuracy in fraud detection and customer behavioral analysis.

The Evolution of Payment Systems: Beyond Rule-Based Models

The world of digital transactions has undergone a massive expansion, with billions of events flowing through systems daily via credit cards, QR codes, contactless payments, and cross-border transfers. This explosion in volume has been matched by an increase in the complexity of financial crime. Fraudsters now leverage generative AI and chatbots to simulate synthetic identities and execute complex, multi-stage attacks.

Historically, payment systems relied on rules-based engines or traditional machine learning models (such as Gradient Boosted Trees) that analyzed data in a “flat” or non-sequential manner. While effective for basic anomalies, these systems often fail to resolve the deep contextual history of a customer. They may miss the subtle shift in behavior that signals a compromised account because they lack the “memory” to connect transactions across long periods. The industry’s challenge is to find a middle way: leveraging the cutting-edge innovation of deep learning while maintaining the explainability and governance required by global financial regulators.

Transformers for Tabular and Sequential Financial Data

The primary innovation discussed is the application of the transformer architecture—originally designed for Natural Language Processing (NLP)—to tabular financial data. Transformers introduce the “attention” mechanism, which allows a model to weigh the importance of different parts of a transaction sequence differently.

In a financial context, this means the model can distinguish between a user’s stable, long-term habits and their recent, potentially anomalous interests. For instance, if a customer who has lived in the same city for ten years suddenly makes a high-value purchase in a foreign country, a transformer can analyze the sequence leading up to that event—looking for “warm-up” transactions or patterns indicative of travel—rather than just flagging the high dollar amount.

Key technical advantages include:

  • Contextual Understanding: Transformers treat the entire transaction history of an entity (customer, merchant, or card) as a sequence, similar to a sentence in a language model.
  • Solving Vanishing Gradients: Unlike Recurrent Neural Networks (RNNs), transformers can capture long-range dependencies without the performance degradation typically associated with long sequences.
  • Multi-Modal Integration: They can blend different data “worlds”—such as event logs, clickstream data, and structured transaction records—into a single global embedding that provides a 360-degree view of an entity.

NVIDIA Accelerated Computing in Financial AI Factories

The training and deployment of these large-scale foundation models require immense computational power, a concept referred to as the “AI Factory.” NVIDIA’s accelerated computing platform is the engine behind these factories, providing the necessary throughput for processing millions of transactions in real time.

NVIDIA’s contribution extends beyond hardware (GPUs like the H100 and Blackwell) to specialized software frameworks. For example, the use of the NVIDIA AI Enterprise suite on AWS allows for efficient tuning and scaling of these models. Furthermore, the integration of Graph Neural Networks (GNNs) with transformers allows systems to not only understand the sequence of transactions but also the relationships between different entities (e.g., shared IP addresses or common merchants among fraudulent accounts). This combined approach enables “pattern mining” at a scale previously thought impossible.

Code Sample: Conceptual Transformer Layer for Transaction Sequences

import torch
import torch.nn as nn

class TransactionTransformer(nn.Module):
    def __init__(self, input_dim, embed_dim, num_heads, num_layers):
        super(TransactionTransformer, self).__init__()
        '''Project tabular transaction features into an embedding space'''
        self.embedding = nn.Linear(input_dim, embed_dim)

        '''Transformer Encoder Layer to capture sequential dependencies'''
        encoder_layer = nn.TransformerEncoderLayer(d_model=embed_dim, nhead=num_heads)
        self.transformer = nn.TransformerEncoder(encoder_layer, num_layers=num_layers)

        '''Output layer for fraud classification (binary: 0 or 1)'''
        self.classifier = nn.Linear(embed_dim, 1)

    def forward(self, x):
        '''# x shape: [batch_size, sequence_length, input_dim]'''
        x = self.embedding(x)
        x = x.permute(1, 0, 2) # Transformer expects [seq_len, batch, embed]
        output = self.transformer(x)
        logits = self.classifier(output[-1]) # Use the last transaction's context
        return torch.sigmoid(logits)

print("Financial Transformer initialized for sequential analysis.")

Real-Time Fraud Detection and Personalized Banking

The ultimate goal of deploying these models on AWS is to move from reactive fraud detection to proactive prevention and hyper-personalization. By leveraging Amazon SageMaker, financial institutions can run “target experiments” and deploy models into a secure, scalable production environment.

The business impact is multifaceted:

  1. Reduced False Positives: By understanding context, models can reduce the number of legitimate transactions being declined, improving customer satisfaction.
  2. Authorization and Routing Optimization: Real-time insights allow for smarter routing of transactions through payment networks, reducing costs and increasing success rates.
  3. Hyper-Personalization: Beyond fraud, these models understand customer intent, allowing banks to offer relevant products and services at the precise moment of need.

While it is still early in the adoption cycle, initial experiments show performance improvements in the range of 1% to 2% in fraud detection accuracy—a seemingly small number that translates into billions of dollars in saved revenue across the global economy.

Conclusion

The intersection of transformer architectures, NVIDIA’s accelerated computing, and AWS’s scalable infrastructure is redefining what is possible in financial services. By treating transaction data as a language to be understood rather than a set of rows to be filtered, the industry is building a more secure and personalized future for global payments. As these “global embeddings” continue to evolve, they will ultimately provide a comprehensive context for every customer, product, and entity in the financial ecosystem.

Links:

PostHeaderIcon [reClojure2025] LLMs + Clojure = Who needs frameworks?

Lecturer

Kapil Reddy is a software engineer known for his “business-first” approach to development. He is a prominent figure in the Clojure community, frequently contributing to discussions and ideation at the Scicloj meetups. Kapil has collaborated with other leading engineers in the ecosystem, such as Vedang Manerikar and Daniel Slutzky, to explore the intersection of artificial intelligence and functional programming. He is currently involved in developing the llms.edn project, which aims to bridge the gap between Clojure’s library-centric philosophy and the modern need for rapid project scaffolding using Large Language Models (LLMs).

Abstract

In the modern software development landscape, Large Language Models (LLMs) have significantly altered workflows, particularly in the realm of project scaffolding. However, the Clojure ecosystem, which prioritizes a philosophy of composable libraries over rigid frameworks, often presents a steep learning curve for newcomers who seek the convenience of “Rails-like” frameworks. This article explores a novel methodology introduced by Kapil Reddy that leverages LLMs to automate the composition of Clojure libraries. By utilizing a structured, native format called llms.edn, developers can describe library usage patterns in a way that LLMs can understand and execute. This approach aims to provide the convenience of a framework while maintaining the flexibility and power of Clojure’s traditional library-based architecture.

The Framework Paradox in Clojure

The debate between using frameworks versus a collection of libraries is central to Clojure’s identity. Traditional frameworks like Ruby on Rails provide a “Golden Path,” offering a set of pre-configured tools and conventions that allow for rapid prototyping. For many developers, especially those transitioning from other ecosystems, the absence of such a framework in Clojure is perceived as a significant barrier to entry. Clojure’s core philosophy leans heavily toward composition, where developers select specialized libraries—such as Ring for HTTP, Reitit for routing, and HugSQL for database access—and manually integrate them.
While this library-centric approach prevents the “black box” complexity and “magic” often associated with frameworks, it requires a deep understanding of the ecosystem. Kapil Reddy observes that LLMs are exceptionally proficient at project scaffolding, a task traditionally reserved for frameworks. The challenge, therefore, is to create a system where LLMs can assist in this scaffolding process without forcing the community to adopt a monolithic framework that would sacrifice the language’s fundamental strengths.

llms.edn: Structured Knowledge for AI Agents

To enable LLMs to effectively compose Clojure libraries, Kapil proposes a structured, Clojure-native approach to describing libraries and their common usage patterns: llms.edn. This concept is inspired by the broader llms.txt initiative but is tailored specifically for the unique requirements of the Clojure ecosystem.
The llms.edn file serves as a manifest that provides the LLM with the necessary context to understand how a library should be initialized, configured, and integrated with others. Instead of the LLM relying on potentially outdated or hallucinatory training data, llms.edn provides a source of truth directly from the library authors or the community. This structured data includes:
* Dependency declarations: Specific coordinates for tools like deps.edn or Leiningen.
* Code snippets: Standard boilerplate for starting a server or connecting to a database.
* Interoperability rules: Instructions on how a library (e.g., a router) interacts with another (e.g., a handler).
By providing these instructions in a machine-readable format, the manual task of “wiring” libraries together—often the most frustrating part for beginners—can be offloaded to an AI agent.

LLM-Powered Composition Workflows

The practical application of this methodology is an LLM-powered composition workflow. In this model, the developer describes the desired features of their application in natural language. An AI agent then queries a registry of llms.edn files to identify the best libraries for the task.
Kapil demonstrates that once the “how-to” for each library is codified, the process of generating a cohesive starter project becomes a “looper making a REST call”. This flow engineering treats the LLM as a pipeline that manages state and passes configuration data between different execution steps. This results in a “framework-like” experience where a full project structure is generated instantly, yet the underlying code remains a collection of simple, independent libraries that the developer can easily modify or replace.
The implications of this shift are profound. It suggests that the primary utility of a framework—reducing the cognitive load of setup and configuration—can now be achieved through intelligent automation. As Kapil notes, the LLM world requires more “simple software” because the models themselves introduce enough complexity; Clojure’s inherent simplicity makes it an ideal target for this kind of AI-driven orchestration.

Links:

PostHeaderIcon [PyDataGlobal2025] What’s Next in AI for Data and Data Management

Lecturer

Lisa Amini is a Distinguished Engineer at IBM and Director of Data & AI Platforms Research, where she also leads IBM’s AI Horizons Network. Her career at IBM Research spans more than two decades and includes foundational work on stream processing systems that became the InfoSphere Streams product, leadership of the IBM Research laboratory in Ireland, and earlier roles directing knowledge and reasoning research. She has guided interdisciplinary efforts across cloud computing, artificial intelligence, and quantum computing, always with an emphasis on technologies that can be deployed at enterprise scale.

Abstract

Recent advances in large language models have catalyzed a wave of AI-assisted tools for data management and operations, ranging from code-generation assistants for data-flow pipelines to retrieval-augmented generation systems and increasingly autonomous data agents. This keynote examines the rapid evolution of generative and agentic capabilities, situates them within the broader data-management stack, and explores both near-term practical applications and longer-horizon research directions. Particular attention is given to the shift from human-operated systems augmented by copilots toward semi-autonomous stacks in which agents design, optimize, remediate, and continuously evaluate data products. The discussion balances technical opportunity with the enduring requirements of price-performance, open-source interoperability, and hybrid data architectures.

The Accelerating Capability Curve and the Emergence of Agency

The pace at which machine-learning benchmarks reach human-level performance has changed dramatically. Tasks that once required decades of incremental progress—handwriting recognition, for example—now reach parity within a few years or even months. Reading comprehension and predictive reasoning benchmarks follow similarly steep trajectories. While these evaluations remain narrow and do not constitute artificial general intelligence, they illustrate an unprecedented rate of improvement. Simultaneously, the cost per inference continues to fall even as model size and training compute grow, a trend driven by better systems design and algorithmic efficiency.

Within this landscape the progression from predictive models to generative models to conversational systems and finally to agents marks a qualitative shift. Agents do not merely answer questions; they dynamically control application flow, make decisions, take actions, and attempt self-correction. In the data domain this agency opens the possibility of systems that no longer wait for humans to formulate every query or repair every broken pipeline. Instead, agents can probe schema, resolve ambiguity, hypothesize data products, evaluate their own output, and iterate.

Transforming the Data Landscape and the Complementary Task Stack

Unstructured data has long existed, yet only recently has it assumed central importance. Machines can now reason over images, generate multimodal content, and extract structured signals from free text at scale. Classical database, warehouse, and lakehouse architectures, optimized primarily for structured tables, must therefore accommodate new access patterns. Retrieval-augmented generation pipelines replace static queries with dynamic retrieval-plus-generation cycles. User interaction moves from fixed application-generated SQL toward speculative, multi-step agent dialogues that probe metadata, formulate candidate queries, and refine them in light of intermediate results.

A useful conceptual inversion is to view the traditional storage–compute–query stack alongside a complementary human-task stack: infrastructure design, workload optimization, data discovery, enrichment, flow creation, remediation, governance, and insight generation. Each of these human activities constitutes fertile ground for agentic automation. Early systems already demonstrate learnable components inside query optimizers and routers; more ambitious research explores whether agents can search the design space of kernel-level software itself.

From Automation to Autonomy: Data Products and Continuous Evaluation

The practical goal is not merely to accelerate individual steps but to move entire workflows from human-operated to human-supervised. Consider the request to stand up a data stack and associated data products for a new application—robo-trading, for instance, that must combine public market data with sentiment signals and support periodic rebalancing. A multi-agent system can be tasked with discovering relevant sources, hypothesizing an ideal schema, populating that schema from heterogeneous tables and documents, extracting structured fields from natural-language text, and packaging the result as a governed data product.

Critical to autonomy is the ability to evaluate quality without constant human intervention. One effective strategy generates natural-language questions that a domain expert would plausibly ask of the intended data product, translates those questions into executable queries, and then monitors coverage metrics (tables and columns touched), topic coverage, query complexity, and latency. Agents iterate—adding sources, refining transformations, simplifying views—until the metrics stabilize within acceptable bounds or progress plateaus and human guidance is required. The same loop can later serve as continuous monitoring: questions that once succeeded can be re-executed to detect drift or regression.

Similar patterns apply to operational remediation. When a data-flow pipeline fails, agents can examine logs, generate natural-language root-cause hypotheses, propose script repairs, and, under appropriate guardrails, test those repairs in a sandbox before presenting them for approval. Across the spectrum of use, build, and optimize activities, the user’s role gradually shifts from operator to approver or observer.

Enduring Constraints and the Research Horizon

Price-performance remains non-negotiable; open-source components continue to enable rapid composition of storage formats, query engines, and table formats; hybrid architectures that span on-premises, cloud, and edge locations persist. Benchmarks, data contracts, open lineage standards, and carefully scoped open-weight models supply the interfaces and evaluation harnesses that allow agents to interoperate safely. Research prototypes already explore operator libraries that let developers request high-level transformations while large language models synthesize the concrete implementations behind the scenes.

The path forward is incremental. Fully autonomous data stacks will not appear overnight. Yet the combination of generative models, agent frameworks, and rigorous evaluation loops is already moving concrete workloads—data-product curation, flow repair, insight generation—along the continuum from assistance toward autonomy. The opportunity for data scientists and engineers is to shape the metrics, tools, and governance practices that will keep these systems both powerful and trustworthy.

Links:

PostHeaderIcon [DevoxxUK2026] Aspiring Speakers: From Replacement to Rocket Fuel – Launching Your Tech Career

Lecturer

Sudi Mandyam is an Engineering Manager at Intradiem, bringing extensive experience in software engineering, site reliability engineering, and cloud technologies. With a background from Visvesvaraya Technological University and roles at organizations including Fastute.io and Navro, Sudi has established himself as a problem solver, leader, writer, and mentor in the technology sector. His insights into AI-driven transformations stem from hands-on leadership in engineering teams navigating rapid industry shifts.

Abstract

In this insightful presentation, Sudi Mandyam challenges prevailing narratives around artificial intelligence displacing developers. Instead, he positions AI as a powerful accelerator for career advancement, particularly for aspiring technologists. Through historical context, evolving AI capabilities, and practical demonstrations, the talk equips attendees with strategies to transition from fearing obsolescence to embracing architectural leadership in an agentic AI era.

The AI Shift: Perception Versus Reality

Sudi opens by highlighting the interconnected nature of technology, opportunities, and problems. He notes that while some perceive AI as a threat to coding professions, this view represents only one facet of a multifaceted evolution. Drawing an analogy to brick-making, he emphasizes that even as AI generates code, human architects remain essential for designing and constructing robust systems.

The presentation traces the rapid progression of AI frameworks over recent years. In 2022, tools like ChatGPT emerged as disruptors, initially seen as potential replacements for search engines. By 2024, solutions such as GitHub Copilot and advanced prompting techniques focused on enhancing speed and efficiency in code generation. However, challenges persisted, including model hallucinations arising from suboptimal prompts or model selections.

Advancing into 2025, agentic programming gained prominence with tools like Cursor and Windsurf, offering improved context handling for microservices and classes, thereby reducing “slop code.” Despite these advances, widespread adoption without adequate guardrails led to security concerns and operational issues. Sudi identifies the current landscape as the “agentic engineering era,” a new discipline layered atop traditional software engineering. Here, context-aware agents function as collaborative colleagues rather than mere coding engines, empowered by frameworks such as CrewAI and Google ADK.

A persistent limitation remains: agents perform only as effectively as the context provided. “Garbage in, garbage out” continues to apply, underscoring the need for sophisticated knowledge management.

Building Organizational Intelligence: LLM Wiki and Intelligent Triage

To address contextual gaps, Sudi introduces the LLM Wiki pattern, inspired by concepts from Andre Karpathy. This approach curates organizational information into a consumable markdown format via an incremental wiki compiler, creating a “second brain” that persists beyond individual experts. Unlike traditional retrieval-augmented generation that may require repeated parsing, the wiki maintains coherent, evolving knowledge repositories.

This second brain proves invaluable across scenarios, particularly incident management. Sudi presents the Intelligent Triage Mesh, which integrates LLM Wiki data, metrics, runbooks, and observability traces from tools like OpenTelemetry and DataDog. A multi-agent orchestration engine evaluates incidents, using confidence thresholds to determine whether automated remediation suffices or human intervention is required.

A live demonstration illustrates these principles in action. Simulating payment failures, an orchestrator leveraging the LLM Wiki decides between auto-remediation and human escalation. Implemented in Go with Google ADK, the system features a main Gemini-powered orchestrator alongside local models for specialized agents. Global policy overrides, managed via the second brain, allow non-technical stakeholders like product managers to update behaviors without code changes.

This methodology significantly improves key metrics such as Mean Time to Recovery (MTTR) within DORA frameworks, transforming incident resolution from hours to minutes.

Conclusion

Sudi Mandyam masterfully reframes AI not as a replacement engine but as rocket fuel for technical careers. By advocating a shift to agentic engineering mindsets and demonstrating practical implementations like contextual wikis and intelligent orchestration, the talk provides actionable pathways for developers to thrive amid technological disruption. Ultimately, the message resonates clearly: problems breed opportunities, and proactive engagement with AI tools positions aspiring speakers and engineers for sustained success.

Links:

PostHeaderIcon [AWSReInvent2025] Supercharging DevOps with AI-Driven Observability: The Next Frontier in SRE

Lecturer

Elizabeth Fuentes is a Senior Developer Advocate at Amazon Web Services (AWS), specializing in the intersection of Artificial Intelligence and DevOps practices. With extensive experience in cloud architecture and software engineering, Elizabeth focuses on how Generative AI can streamline complex CI/CD pipelines and enhance Site Reliability Engineering (SRE). She is a key contributor to AWS educational initiatives, having co-developed advanced courses on AI-driven automation. Joining her is Laas Alina, a software architect and open-source enthusiast who focuses on implementing multi-agent systems and the Model Context Protocol (MCP) to solve observability challenges at scale.

Abstract

As software systems grow increasingly distributed and complex, traditional observability—centered on manual log analysis and reactive dashboards—is becoming insufficient. This article explores the paradigm shift toward AI-driven observability, where Generative AI serves not just as a query tool, but as an active participant in failure detection, correlation, and resolution. By leveraging Amazon Bedrock and Amazon Q, organizations can transition from “reactive” to “predictive” DevOps. The discussion analyzes the methodology of building AI agents that simulate architectural stress, automatically explain multi-layered failures, and provide traceable, actionable recommendations. We examine the implementation of the Model Context Protocol (MCP) in establishing sophisticated multi-agent systems (MAS) that transform raw data into contextual understanding, ultimately reducing the Mean Time to Resolution (MTTR) and enhancing systemic resilience.

The Evolution of Observability: From Metrics to Contextual Understanding

The traditional pillars of observability—metrics, logs, and traces—provide the “what” of a system’s state but often fail to provide the “why” in real-time. In high-velocity DevOps environments, the sheer volume of telemetry data can overwhelm human operators, leading to “alert fatigue” and delayed responses to critical incidents. Elizabeth posits that the integration of Generative AI marks the fourth pillar of observability: Contextual Intelligence. This evolution moves the industry beyond simple threshold-based monitoring toward systems that understand the semantic relationship between a failed deployment, a spike in latency, and a specific line of code.

By utilizing Large Language Models (LLMs) through Amazon Bedrock, DevOps teams can ingest vast amounts of unstructured log data and receive summaries that highlight anomalies that might be missed by traditional regex-based filters. The methodology involves training the AI to recognize “normal” operational patterns and identifying deviations not just by value, but by the intent of the system’s behavior. This contextual layer allows for a more nuanced interpretation of system health, where the AI can distinguish between a benign resource spike and a precursor to a cascading failure.

Architecting AI Agents for Predictive Troubleshooting

The transition to AI-driven observability is characterized by the deployment of “Micro-agents”—specialized AI entities designed to handle specific segments of the DevOps lifecycle. These agents operate within a Multi-Agent System (MAS), where they collaborate to solve complex incidents. For instance, a “Monitoring Agent” might detect a performance degradation and immediately trigger a “Diagnosis Agent” to correlate the event with recent CI/CD pipeline changes.

Elizabeth and Laas Alina emphasize the importance of the Model Context Protocol (MCP) in this architecture. MCP acts as the communication backbone, allowing agents to share context without losing the “lineage” of a decision. When an AI agent recommends a specific architectural change or a rollback, it must provide clear traceability. This is crucial for maintaining trust in automated systems. The agents do not operate in a vacuum; they interact with tools like Amazon Q to provide developers with instant explanations of failures directly within their Integrated Development Environment (IDE) or chat interface.

// Example of an AI-driven Observability Agent Configuration
agent:
  name: "IncidentDiagnosticAgent"
  provider: "AmazonBedrock"
  model: "claude-3-sonnet"
  capabilities:
    - log_analysis
    - metric_correlation
    - trace_summarization
  mcp_config:
    protocol_version: "1.0"
    shared_context: "deployment_metadata"
  safety_guardrails:
    - max_token_usage: 4000
    - human_in_the_loop_required: true

Transforming CI/CD through Generative AI and Simulation

Beyond reactive troubleshooting, AI-driven observability empowers proactive system design. One of the most innovative concepts discussed is the use of AI agents to simulate “stress-test” scenarios within a digital twin of the production environment. These agents can intentionally inject failures—similar to Chaos Engineering—and then observe how the observability stack responds. This creates a feedback loop where the AI helps engineers identify “blind spots” in their monitoring before a real incident occurs.

Furthermore, Generative AI transforms the CI/CD pipeline by automatically generating “failure explanations.” Instead of a developer sifting through a 5,000-line build log, Amazon Q can provide a concise summary: “The build failed because the new database schema in commit X is incompatible with the connection pool settings in environment Y.” This level of automated insight accelerates the “inner loop” of development, allowing engineers to focus on innovation rather than infrastructure archeology.

The Human-AI Partnership: Strategic Implications

A common concern in the industry is the replacement of human engineers by AI. However, Elizabeth argues that the future belongs to the “augmented engineer.” AI is a force multiplier that automates the repetitive, “drudge work” of observability—log parsing and initial triage—allowing human experts to focus on high-level strategy and complex architectural decisions. The goal is to transform teams from being “reactive” (fighting fires) to “proactive” (preventing fires).

Implementing these systems requires a cultural shift toward AI-literacy within DevOps teams. Organizations must establish safety guardrails to ensure that AI-driven recommendations are validated and that automated actions (like auto-remediation) have clear rollback paths. By embracing AI as a strategic tool, DevOps and SRE teams can achieve a level of operational excellence that was previously unattainable, ensuring that as systems grow in scale, their reliability grows in parallel.

Links:

PostHeaderIcon [DevoxxGR2026] Code That Moves the World: The Rise of Physical AI

Lecturer
Will Sentance is the founder of Standard Material and Codesmith, organizations at the forefront of physical AI infrastructure and AI/software engineering education. A speaker, educator, and practitioner, Sentance bridges software engineering expertise with emerging robotics and autonomous systems. He contributes to research at Oxford and leads initiatives training talent for the next wave of intelligent physical systems.

Abstract
In this forward-looking keynote at Devoxx Greece 2026, Will Sentance explores the profound convergence of software engineering and physical intelligence. Robots and autonomous systems are transitioning from specialized, brittle demonstrations to capable, generalizable agents operating in real-world environments. Sentance details the technological breakthroughs in hardware, data, and foundation models driving this transformation and argues that traditional software engineering skills are central to building the platforms, data pipelines, and integrations required for scalable physical AI deployment.

The Remarkable Progress in Physical Intelligence

Physical AI—systems that sense, understand, and act upon the physical world—has advanced dramatically. Robots now follow natural language instructions, handle novel objects, and demonstrate emergent capabilities. Foundation models for robotics enable zero-shot generalization and long-horizon planning across diverse embodiments.

Companies like Physical Intelligence, Agility Robotics, and others are moving from laboratory experiments to industrial and domestic applications. This shift is fueled by massive investment and rapid iteration.

Core Technological Enablers

Three key areas have transformed the landscape:

Hardware Revolution: Affordable, off-the-shelf components—from full humanoids to grippers and sensors—dramatically lower barriers. Edge computing platforms provide sufficient power for onboard inference.

Data Explosion: Teleoperation, simulation (including sophisticated world models), and real-world deployment generate multimodal datasets at unprecedented scale. Techniques like action chunking address real-time requirements.

AI Models: End-to-end learning replaces traditional control theory. Vision-language-action models predict continuous action trajectories, enabling flexible behavior without exhaustive manual programming.

The Physical AI Technology Stack

Sentance outlines a layered architecture:

  • Real-time Control: Low-level, deterministic operations managing actuators and safety at high frequency.
  • Platform and Middleware: Abstractions like ROS providing integration, simulation interfaces, and developer tools.
  • Intelligence Layer: Foundation models processing vision, language, and proprioception to generate actions.
  • Data and Learning Loop: Continuous collection, training, evaluation, and deployment cycle.

Opportunities for Software Engineers

Contrary to initial impressions, software engineers are perfectly positioned to lead this revolution. Approximately 80% of the required work involves familiar disciplines: systems architecture, platform engineering, data pipelines, low-level optimization, and agentic integration.

Roles at leading organizations emphasize scalable frameworks, reliable deployment, observability, and integration of AI models into production—skills honed in cloud-native and distributed systems development.

New challenges center on real-time constraints, physical dynamics, and managing massive multimodal datasets, but these build directly upon existing expertise.

Getting Started with Physical AI

Sentance encourages practical experimentation using affordable hardware like the SO-101 and open tools. Developers can quickly train policies for simple tasks such as closing a laptop lid, experiencing the full cycle from data collection to deployment.

The physical world represents the next major platform for code. Software engineers who embrace this frontier will shape the coming industrial transformation.

Links: