Recent Posts
Archives

Posts Tagged ‘DeveloperExperience’

PostHeaderIcon [DevoxxGR2026] The Pragmatic Path: Structured Adoption of Agentic AI in Software Development

Lecturers
Dimitris Papageorgiou and Konstantina Mavrodimitraki are Senior Solutions Architects at Amazon Web Services in Greece. With extensive experience as software and data engineers, they have supported numerous enterprise customers in adopting cloud-native and AI technologies. Their work focuses on practical implementation strategies that deliver measurable business value while addressing real-world concerns around quality, security, and team readiness.

Abstract
Dimitris Papageorgiou and Konstantina Mavrodimitraki present a pragmatic framework for integrating agentic AI into software development lifecycles. Based on hands-on implementations across multiple customer environments, the session addresses common barriers such as code quality fears, lack of structure, and resistance to change. Through concrete examples—including optimized code reviews returning over 16,000 developer hours annually and 65-80% faster issue resolution—they outline a phased approach from individual experimentation to cross-team standardization and organizational scaling.

The Current State of AI Adoption in Development Teams

Many organizations purchase AI tool licenses and distribute them broadly, expecting immediate productivity gains. In practice, developers experiment individually—often engaging in “vibe coding”—without shared practices or metrics. This leads to fragmented adoption, inconsistent quality, and difficulty demonstrating return on investment to leadership.

The speakers identify a critical gap: while tools proliferate, teams lack a common language and structured methodology. Success requires moving beyond ad-hoc usage to deliberate integration aligned with specific pain points.

A Framework for Systematic Agentic AI Adoption

The proposed framework operates along two dimensions: organizational pain points and AI maturity levels. Pain points—such as code review bottlenecks, testing coverage, or feature development velocity—must be identified first. Maturity progresses from individual experimentation to team standardization and finally cross-team integration.

Teams begin at their current maturity level and implement solutions appropriate to that stage. For code review bottlenecks, level-one teams conduct structured experimentation with various tools, followed by retrospectives to select winners. Level-two teams document guidelines, define success metrics, and establish processes. Level-three organizations embed AI into pipelines with shared patterns and governance.

Applying the Framework: Code Reviews and Testing

For code reviews, a real-world AWS customer in betting and gaming implemented an agentic workflow using Amazon Bedrock. Pull request events trigger enrichment via data pipelines before an agent analyzes changes against coding standards, security rules, and business requirements. The system posts comments directly, with optional human validation.

Metrics showed over 16,700 developer hours returned annually, allowing focus on higher-value work. Similar patterns apply to testing: starting with AI-assisted unit test generation, teams progress to standardized pipelines and shared test patterns across the organization.

Feature Development with Spec-Driven Approaches

Spec-driven development extends AI assistance across the lifecycle. Rather than isolated prompts, teams collaborate with agents to refine requirements, architectural decisions, and task breakdowns. Amazon Q Developer exemplifies this, generating user stories, acceptance criteria, designs, and implementation tasks from high-level intents.

This approach reduces back-and-forth during sprint planning and ensures generated code aligns with broader context. Workshops help teams adapt the process to their needs, fostering ownership and continuous improvement.

Scaling and Avoiding Common Pitfalls

Successful scaling requires executive sponsorship, dedicated time for experimentation, and clear metrics. Leadership must treat AI adoption as a strategic initiative rather than a side project. Engineers should share learnings and metrics to build momentum.

Pitfalls include unstructured experimentation leading to technical debt, over-reliance on AI without human oversight, and failure to measure impact. The speakers recommend divide-and-conquer: tackle one pain point thoroughly before expanding.

Conclusion: AI as a Multiplier of Good Practices

Agentic AI amplifies existing strengths in clean code, testing, documentation, and collaboration. By following a pragmatic, maturity-aligned path, teams achieve faster delivery, higher quality, and greater developer satisfaction. The framework transforms AI from a hype-driven experiment into a structured capability delivering tangible results.

Links:

PostHeaderIcon [GopherConUK2025] A Gopher’s Guide to Vibe Coding: Evaluating LLM-Assisted Software Engineering in Go

Lecturer

Daniela Petruzalek works as a Developer Relations Engineer at Google. Originally from Brazil and residing in the United Kingdom since 2019, Daniela has worked extensively with the Go programming language since 2017. Her professional background spans software development, system architecture, agile technical consulting, and developer advocacy, with a primary focus on cloud-native technologies, developer experience, and language tooling.

Abstract

The rapid evolution of Large Language Models (LLMs) has popularized “vibe coding”—a development paradigm where software engineers rely on natural language prompts to drive automated code generation. While early iterations centered on unchecked code synthesis, professional adoption requires incorporating LLM tools into structured engineering workflows. This paper examines the integration of LLM coding agents into idiomatic Go development, evaluating productivity, correctness, and code quality. Drawing from practical implementations—including the creation of testquery and Model Context Protocol (MCP) servers—this study details techniques such as context engineering, Retrieval-Augmented Generation (RAG), tool-based grounding, and automated peer-review loops to produce maintainable, production-ready Go code.

Taxonomy of AI-Assisted Development Tools

AI-assisted software engineering tools range from simple inline completion utilities to fully autonomous agents:

+-------------------------------------------------------------------+
|               AI ASSISTED DEVELOPMENT SPECTRUM                    |
+-------------------------------------------------------------------+
|  Inline Completion  |  Contextual Chat  |  CLI Agents  | Autonomous|
|  (GitHub Copilot)   |  (VS Code Chat)   | (Gemini CLI) |  (Jules)  |
+-------------------------------------------------------------------+
  Low Autonomy <-------------------------------------> High Autonomy

  1. Inline Code Completion: Algorithms that complete single lines or block structures using local code context.
  2. Contextual Chat Interfaces: In-IDE conversational assistants capable of querying localized workspace snippets.
  3. CLI-Based Coding Agents: Interactive command-line tools (e.g., Gemini CLI, Aider) capable of running local shell commands, inspecting file systems, executing tests, and applying multi-file edits directly.
  4. Autonomous Execution Agents: Fully asynchronous environments (e.g., Jules, Devin) that clone repositories within isolated virtual machines, resolve GitHub issues, execute test suites, and submit pull requests with minimal human intervention.
// Sample Model Context Protocol (MCP) tool declaration in Go
package main

import (
    "context"
    "fmt"
)

type GoDocRequest struct {
    Package string `json:"package"`
    Symbol  string `json:"symbol,omitempty"`
}

func HandleGoDoc(ctx context.Context, req GoDocRequest) (string, error) {
    if req.Package == "" {
        return "", fmt.Errorf("package name is required")
    }
    // Tool execution logic returning package documentation
    return fmt.Sprintf("Documentation for %s", req.Package), nil
}

Integrating autonomous agents into professional workflows requires a structured framework based on business value and technical certainty. High-value tasks with clear technical steps demand real-time human oversight via interactive CLI tools. Conversely, low-risk, repetitive tasks—such as updating open-source license headers or formatting readmes—can be delegated to asynchronous background agents.

                Business Value vs Technical Certainty
               +-------------------+-------------------+
               | Research / Spikes |   Top Priority    |
  High Value   | (Deep Research /  | (Synchronous CLI  |
               |  Hands-on-Keys)   |   Development)    |
               +-------------------+-------------------+
               |    Not Doing      |   Nice-to-Have    |
  Low Value    | (Ignore / Backlog)|  (Asynchronous    |
               |                   | Autonomous Agent) |
               +-------------------+-------------------+
                   Low Certainty       High Certainty

Context Engineering, Grounding, and Context Degradation

Generative models encounter several structural limitations when handling source code, including out-of-date training data, non-deterministic outputs, and a tendency to hallucinate invalid package APIs. Overcoming these limitations requires precise context engineering and tool grounding.

  • Context Engineering: Supplying targeted technical documentation directly within the prompt scope. Fetching fresh package definitions prevents models from using deprecated function signatures or obsolete import paths.
  • Tool Grounding: Providing external capabilities—such as file system readers, web scrapers, or language server protocols—via structured mechanisms like the Model Context Protocol (MCP). Grounding allows LLMs to query documentation dynamically rather than relying solely on parametric memory.
// Example Go code illustrating explicit error handling for LLM tasks
package main

import (
    "errors"
    "fmt"
)

var ErrInvalidSymbol = errors.New("requested symbol not found")

func LookupSymbol(doc, symbol string) (string, error) {
    if symbol == "" {
        return doc, nil
    }
    // Explicit string processing logic
    return "", ErrInvalidSymbol
}

Prolonged agent sessions suffer from context degradation (or context rot), where accumulated error logs, discarded attempts, and verbose outputs pollute the model’s working memory. This degradation can cause the model to repeat failed edits or enter infinite refactoring loops. Engineers must actively manage context length by resetting sessions (/clear), providing fresh context, or summarizing state transitions before continuing development.

The Iterative TDD Refactoring Loop and Multi-Agent Code Reviews

Unchecked vibe coding often results in fragile, unmaintainable implementations. To ensure production-grade software quality, vibe coding should follow the disciplined Test-Driven Development (TDD) cycle:

      +-------------------------------------------------+
      |                                                 |
      v                                                 |
+-----------+        +-----------+        +-----------+ |
| RED Phase | -----> | GREEN     | -----> | REFACTOR  | -+
| Write Test|        | Pass Test |        | Code Review|
+-----------+        +-----------+        +-----------+

  1. Red Phase: Define failing tests or precise specification prompts detailing constraints, input structures, and expected outcomes.
  2. Green Phase: Direct the LLM to generate the minimal implementation required to pass the test suite.
  3. Refactor Phase: Enforce code readability, performance optimizations, and idiomatic Go practices before starting new features.
// Example table-driven test to enforce green-phase verification
package main

import "testing"

func TestLookupSymbol(t *testing.T) {
    tests := []struct {
        name    string
        doc     string
        symbol  string
        wantErr bool
    }{
        {"empty symbol", "package doc", "", false},
        {"missing symbol", "package doc", "Foo", true},
    }

    for _, tt := range tests {
        t.Run(tt.name, func(t *testing.T) {
            _, err := LookupSymbol(tt.doc, tt.symbol)
            if (err != nil) != tt.wantErr {
                t.Errorf("LookupSymbol() error = %v, wantErr %v", err, tt.wantErr)
            }
        })
    }
}

A key strategy for maintaining code quality is decoupling implementation from code review. Allowing the same LLM session to review its own output often yields false positives due to lingering context bias. Instead, developers should pass the newly generated code to a fresh, isolated LLM instance equipped with dedicated code-review system prompts. This independent review step effectively surface edge-case bugs, unhandled errors, unreachable code, and style violations.

+----------------------+         +----------------------+
|  Primary Coding Agent |         | Isolated Review Agent |
| (Generates Solution) |         | (Fresh Context Window)|
+----------------------+         +----------------------+
           |                                |
           | Output Source Code             | Analyzes AST & Patterns
           +------------------------------->|
                                            |
                                            v
                                 +----------------------+
                                 |  Structured Feedback |
                                 |  (JSON Diagnostics)  |
                                 +----------------------+

By pairing automated multi-agent code reviews with explicit project guidelines (e.g., AGENTS.md or GEMINI.md), developers create a self-improving feedback loop. As the agent runs reflection prompts to analyze past session mistakes, it systematically updates its instruction rules, raising code quality in subsequent sessions.

Links:

PostHeaderIcon [DotJs2024] API Design is UI Design

In the intricate tapestry of software craftsmanship, the boundaries between visual interfaces and programmatic ones blur, revealing a unified discipline rooted in empathy and usability. Lea Verou, a luminary in web standards and W3C TAG member, delivered a revelatory session at dotJS 2024, asserting that API design mirrors UI design in every facet—from intuitiveness to error resilience. With a PhD from MIT focused on developer experience and stewardship of dozens of open-source projects, Verou dissected the pitfalls of APIs that frustrate and the principles that enchant, urging creators to treat code as an interface wielded by fellow humans.

Verou opened with a visceral anecdote: the SVG DOM’s labyrinthine quest to extract a circle’s radius, yielding not a crisp number but an SVGAnimatedLength riddled with baseVal, animVal, and unit-conversion methods—annoying even sans animations. This exemplar encapsulated her thesis: APIs, be they functions, classes, components, or native browser APIs, are user interfaces where developers are the users, and interactions manifest as keystrokes. Echoing Alan Kay’s maxim—”simple things should be easy, complex things possible”—she mapped it to a complexity plane: low-effort dots for trivial tasks, viable paths for sophistication. Usability tenets, from Google Calendar’s drag-and-drop simplicity to advanced recurrence rules, permeate both realms; DX is merely UX recast for code scribes.

Central to Verou’s discourse was user-centricity: APIs thrive when attuned to genuine needs, not theoretical purity. High-level use cases—like assembling IKEA furniture with a screwdriver—inform broad abstractions, while low-level ones—like screwing a wall anchor—demand primitives. She critiqued legacy DOM traversals burdened by redundant parent references, supplanted by modern APIs favoring single-truth sources. Components exemplify elegance: encapsulating dialog boilerplate into reusable units slashes cognitive load. Iterating isn’t prohibitive; ship high-level facades covering 80% of scenarios, layering primitives as demands surface—or vice versa, observing usage to scaffold abstractions. The Intl.DateTimeFormat API’s evolution—from vague toLocaleString to nuanced options yielding structured outputs—exemplifies this progressive disclosure, smoothing from casual to granular control.

Verou championed empirical validation: user testing, sans visuals, via representative tasks and think-aloud protocols. Five participants unearth 85% of issues; two halve them. Zoom suffices—no labs required. Dogfooding complements: prototype demos, draft docs, author tests pre-implementation, refining iteratively. Empathy crowns all: intuit users’ pains, infer principles organically. Tailwind’s rise signals CSS’s accessibility gaps; blame the medium, not the maker, and mend it. Verou’s clarion call: infuse humanity into APIs, easing burdens and amplifying creativity across the dev spectrum.

Usability Principles Across Interfaces

Verou wove Kay’s dictum into a visual quadrant, plotting task complexity against UI/API effort, advocating coverage of simple-easy and complex-possible quadrants. User needs—pain points, scenarios—drive this: distinguish macro goals from micro actions, ensuring APIs mirror real workflows. SVG’s unit obsessions ignored 90% of queries; streamlined getters would suffice. Progressive layers, as in date formatting’s escalating options, democratize power without overwhelming novices.

Empirical Refinement and Empathy

Testing APIs demands observation: task users, query struggles non-leadingly, affirm it’s the design under scrutiny. Verou debunked myths—engineers aren’t users; widespread misuse indicts the API. Dogfood rigorously: sketch code flows pre-build, iterate via docs and tests. Ultimate imperative: cultivate care—empathize with wielders’ contexts, yielding designs that intuit, adapt, and inspire.

Links: