Posts Tagged ‘LLM’
[SpringIO2026] Hybrid Modernization: Combining OpenRewrite’s Precision with LLM Intelligence for Spring
Lecturer
Raquel Pau is a technical product manager at Broadcom (formerly VMware Tanzu). She brings extensive experience in Java developer tools, continuous-integration and continuous-delivery platforms, and internal developer platforms. Previously she worked as an engineering manager at Moderne, the company behind OpenRewrite, and held product-management roles at CloudBees focused on developer productivity. She has spoken at multiple Spring I/O editions as well as Devoxx, JavaConf and JavaZone. Her background combines deep technical knowledge of code-transformation tooling with product thinking about how large organizations can keep their application portfolios modern and consistent.
Abstract
Code modernization is not a single problem. Upgrading a Spring Boot application within the same major version, migrating from JAX-RS to Spring MVC, and rewriting a COBOL batch job into Spring Batch demand fundamentally different strategies. This article explores the taxonomy of modernization tasks proposed by Raquel Pau and the hybrid methodology that pairs OpenRewrite’s deterministic, type-aware recipes with the semantic reasoning power of large language models. Concrete demonstrations illustrate how upgrade plans are calculated from Maven metadata, how skills orchestrate recipe execution followed by LLM-driven semantic fixes, and how a structured DSL extracted from legacy code guides a full rewrite while preserving contracts and enabling incremental delivery.
Deterministic versus Non-Deterministic Transformations
Modernization tools fall into two broad categories. Deterministic tools always produce the identical output for a given input. Renaming a method, updating a package import, or replacing a deprecated Spring API are deterministic operations. OpenRewrite belongs to this category: it operates on a lossless semantic tree that retains type attribution obtained from the compiler, applies visitor-based recipes, and preserves the original formatting of the source. Because the transformation is deterministic, recipes can be unit-tested with high confidence and executed at scale across hundreds of repositories without surprise.
Non-deterministic problems admit many correct answers. Generating documentation, extracting the business intent of a filter, or inventing an idiomatic Spring Security configuration from a set of JAX-RS name-binding annotations are examples. Large language models excel here because they reason over patterns and can synthesize higher-level constructs that do not exist in the original code. The cost, however, is variability, the need for evaluation harnesses, and a tendency to hallucinate when internal libraries or proprietary APIs are outside the model’s training distribution.
OpenRewrite’s limitations are the mirror image of its strengths. It cannot perform runtime analysis; dependency injection and reflection mean that many object relationships become visible only after the application starts. It cannot invent new semantic abstractions; a mechanical translation of JAX-RS filters into Spring filters often leaves residual compilation errors or suboptimal configurations that require human or LLM insight. Cross-language migration is outside its design scope.
Coding agents partially compensate for these gaps by using pattern-based reasoning and by iterating until the project compiles. Yet they lack default type attribution, suffer from context-window constraints, and generate large volumes of tokens before reaching a stable state. The rational strategy is therefore hybrid: apply deterministic recipes first to shrink the problem, then invoke the LLM only for the residual semantic work.
Three Levels of Modernization
Pau organizes modernization into three progressively more demanding levels.
Upgrades remain inside the same framework family. A Spring Boot 3.3 application is moved to Spring Boot 4, simultaneously updating transitive dependencies such as Jackson and JUnit. Because Spring’s release train is not strictly linear and because organizations maintain internal frameworks with their own release cadences, a simple “latest version” recipe is insufficient. An upgrade-plan engine inspects Maven metadata, calculates a sequence of compatible intermediate steps, and emits a series of small, reviewable pull requests. Each step leaves the application in a buildable state. Tanzu’s Application Advisor exposes this capability via the cf repo upgrade plan and cf repo apply upgrade plan commands, demonstrating that continuous, low-risk upgrades can be embedded in CI pipelines.
Migrations change the underlying framework while preserving language and runtime. The canonical example is Jakarta JAX-RS to Spring Boot. Name-binding annotations that attach filters to resources have no direct counterpart; authentication filters must become Spring Security configurations; repositories must acquire @Repository annotations. The recommended skill therefore first executes the OpenRewrite recipes that perform the mechanical rewrite and any accompanying Spring Boot upgrade, then hands control to the coding agent to resolve remaining compilation errors and to map name-binding semantics onto Spring constructs. The result is both more complete and far less expensive in tokens than asking an unconstrained LLM to rewrite the entire application.
Full rewrites discard the original implementation while preserving contracts. A COBOL batch program that sorts records by date and amount must become a Spring Batch job that reads the same input format, produces identical output, and respects the same database schema if one is involved. Because legacy systems rarely possess comprehensive tests, the process begins by extracting a catalog of user stories, then a structured domain-specific language description of inputs, outputs, and processing steps. Only after the human reviewer validates the generated tests and the semantic model does the agent emit Spring code, typically seeded by a skeleton obtained from start.spring.io. Incremental delivery is essential: large monolithic rewrites cannot be reviewed or risk-managed in a single step.
Orchestrating OpenRewrite and LLM Agents
Three integration mechanisms allow a coding agent to invoke OpenRewrite without saturating its context window. Local MCP servers expose the rewrite CLI so that only the command and its concise output enter the conversation. Skills package the same CLI invocation and are loaded only when the agent decides the skill is relevant. Prompts can be registered with a remote MCP server, yet they must be fully present in every conversation and therefore scale poorly for complex migrations.
The hybrid skill for a JAX-RS migration therefore looks roughly as follows: calculate the upgrade plan that includes the JAX-RS recipes, execute the recipes, collect residual compilation diagnostics, and finally apply semantic transformations that replace name-binding filters with Spring Security and Spring MVC constructs. Because the deterministic phase has already performed the bulk of the mechanical work, the LLM operates on a far smaller residual problem and produces higher-quality results.
For full rewrites the skill is organized into three explicit phases. Phase one extracts a user-story catalog and stores it under version control so that subsequent runs reuse the analysis. Phase two materializes a structured DSL for a chosen story, including acceptance criteria, data models, and external contracts. Phase three generates the Spring implementation and correlating tests. Human validation remains mandatory; the agent cannot be trusted to invent missing requirements or to decide whether an original implementation was correct.
Practical Demonstrations and Organizational Implications
In the upgrade demonstration a Spring Petclinic application on Boot 3.3 is analyzed; the engine proposes coordinated upgrades of Spring Boot, Jackson and JUnit; successive apply steps produce small, reviewable diffs that leave the project green after each commit. In the migration demonstration a pure JAX-RS Petclinic is transformed: OpenRewrite rewrites the bulk of the code, the agent resolves compilation issues caused by signature changes, and name-binding annotations disappear in favor of proper Spring Security configuration. In the rewrite demonstration a simple COBOL sorter is analyzed, a single user story and its DSL are generated, a Spring Batch project is scaffolded, and the resulting executable produces byte-for-byte identical output.
The organizational payoff is standardization. When every application can be moved to a common Spring Boot baseline with low friction, teams share libraries, security configurations and operational practices. Token consumption drops dramatically because deterministic recipes eliminate the majority of mechanical work. Evaluation of non-deterministic skills becomes feasible because the residual problem set is smaller and more homogeneous.
Conclusion
Modernization success depends on matching the tool to the nature of the transformation. OpenRewrite supplies precision, testability and scalability for deterministic changes. Large language models supply the semantic insight required for migrations and rewrites. A carefully designed hybrid that keeps the LLM outside the hot path of routine upgrades, that constrains its context to residual problems, and that forces explicit contracts for full rewrites yields both higher quality and lower cost. Organizations that adopt this disciplined approach can keep large application portfolios current without sacrificing reviewability or operational safety.
Links:
[DevoxxFR2026] Measuring the Unmeasurable: Evaluating Generative AI Systems
Lecturer
Erin Pacquetet is an expert in AI evaluation and product development at SCIAM, a Paris-based consulting firm. With a background in linguistics and extensive experience guiding enterprises through the complexities of deploying generative AI applications, she specializes in bridging technical implementation with business requirements and robust quality assurance.
Abstract
Generative AI systems promise transformative capabilities but present unique evaluation challenges due to their creative and unpredictable nature. Erin Pacquetet addresses this paradox by outlining comprehensive strategies for assessing systems that blend linguistic fluidity with strict factual accuracy. Using a Retrieval-Augmented Generation (RAG) chatbot as a running case study, the presentation examines limitations of traditional metrics, the role of LLM-as-a-judge approaches alongside their inherent biases, the necessity of human evaluation, and continuous monitoring to detect drift. Attendees gain practical frameworks for building reproducible evaluation pipelines that balance innovation with reliability in production environments.
The Fundamental Challenge of Evaluating Generative Systems
Generative AI introduces a core tension between creativity and control. Organizations adopt large language models precisely because they handle diverse, uncontrolled inputs and produce personalized outputs. Yet this very strength complicates evaluation. Traditional deterministic testing works for rule-based systems but falls short when outputs vary naturally while needing to remain accurate, relevant, and safe.
In the case study of an insurance company’s customer-facing RAG chatbot, the system must answer questions about policies while adhering to brand tone, regulatory constraints, and response length limits. A single question like “Is home insurance mandatory for tenants in France?” could yield multiple valid responses of varying quality. Evaluation must therefore move beyond binary correctness to nuanced assessment across multiple dimensions.
Effective evaluation pipelines transform qualitative judgments into quantitative, scalable measurements. This requires simulating realistic inputs, generating outputs, and assessing them against well-defined criteria. The process must cover ideal scenarios, expected real-world usage, and adversarial cases to ensure robustness before production deployment.
Simulating Inputs: Ideal, Realistic, and Adversarial Scenarios
The foundation of any evaluation lies in a carefully constructed dataset representing the full spectrum of potential interactions. For the insurance chatbot, inputs fall into three categories.
Ideal inputs are perfectly formed questions with clear intent and complete context, such as grammatically correct inquiries directly related to covered products. These establish baseline performance and set high acceptance thresholds.
Realistic inputs mirror actual user behavior: keyword-based queries, vague phrasing, oral-style language, partial context, or minor errors. Testing these ensures the system handles the messy reality of production traffic rather than sanitized examples.
Adversarial inputs probe vulnerabilities: prompt injections, attempts to elicit harmful content, off-topic questions, or malicious efforts to bypass safeguards. These reveal security weaknesses and edge cases that could damage reputation or expose risks.
Creating this dataset demands collaboration between technical teams and domain experts. Business stakeholders define what constitutes success for each category, translating abstract requirements into concrete examples. This exercise often reveals inconsistencies in initial specifications, forcing clarification before development advances.
The resulting evaluation dataset serves as both a benchmark and a living artifact. It evolves with the product, incorporating new failure modes discovered in production and expanding coverage as usage patterns emerge.
Generating and Assessing Outputs: Metrics and Human Judgment
Once inputs are prepared, the system generates outputs for evaluation. Assessment occurs along two primary axes: output quality and operational performance.
Output quality encompasses factual accuracy, relevance to the query, completeness of information, and safety. For the RAG chatbot, responses must draw correctly from policy documents, address the specific question asked, provide sufficient detail without excess length, and maintain an appropriate empathetic tone.
Traditional metrics prove insufficient. String matching fails to capture semantic equivalence across varied phrasings. Semantic similarity measures can overlook critical omissions or subtle inaccuracies. Probabilistic approaches, particularly LLM-as-a-judge, offer greater flexibility by leveraging models to analyze outputs against detailed criteria.
A well-crafted judge prompt might instruct the model to identify contradictions or omissions between a generated response and a reference answer, returning a binary judgment. This constrains the evaluation task sufficiently to reduce variance while maintaining nuance. Multiple specialized judges can target different aspects: one for factual consistency, another for tone alignment, and a third for regulatory compliance.
Human evaluation remains essential for validation. Domain experts review samples to calibrate automated metrics, ensuring alignment between machine judgments and business expectations. This human-in-the-loop process establishes confidence thresholds for each metric.
Operational metrics complement quality assessment. Response latency, cost per inference, and system stability must meet production requirements. A perfectly accurate but slow response fails as a product. Monitoring these dimensions alongside quality creates a holistic view of system readiness.
Building and Maintaining Evaluation Pipelines
A complete pipeline integrates input simulation, output generation, and multi-faceted assessment into an automated workflow. Teams execute evaluations frequently: after prompt modifications during development, before major releases, and continuously in production to detect regression or drift.
The evaluation dataset evolves as the central reference point. Production logs reveal new query patterns or failure modes, which teams incorporate to strengthen coverage. Regular human review sessions ensure metrics remain aligned with changing business needs and user expectations.
For the insurance chatbot, this meant balancing completeness against brevity, factual precision against approachable language, and safety against helpfulness. The dataset captured these trade-offs explicitly, allowing systematic optimization rather than guesswork.
Challenges persist. Judge models can inherit biases or exhibit inconsistency. Human evaluators introduce subjectivity. Thresholds require careful tuning to avoid both false confidence and excessive caution. Success demands iterative refinement and cross-functional collaboration.
From Evaluation to Production Confidence
Robust evaluation bridges the gap between promising prototypes and reliable production systems. By systematically addressing the inherent variability of generative outputs, teams build the confidence necessary for deployment.
The insurance chatbot case demonstrates that evaluation is not merely technical validation but a strategic discipline. It forces clarification of requirements, surfaces hidden assumptions, and creates shared understanding across technical and business stakeholders.
As generative AI proliferates, organizations that master evaluation gain competitive advantage. They deploy innovative capabilities with appropriate safeguards, iterating rapidly while maintaining quality. The discipline transforms the “unmeasurable” into something manageable, turning potential risk into sustainable value.
Links:
[VoxxedDaysBucharest2026] Optimizing LLM Inference on Kubernetes: Abdel Sghiouar on Practical Techniques for the Rest of Us
Lecturer
Abdel Sghiouar is a Developer Advocate at Google Cloud with deep expertise in cloud-native technologies, Kubernetes orchestration, and AI/ML workload optimization. Drawing from a robust background in infrastructure engineering and open source contributions, Abdel helps organizations design, deploy, and tune complex AI applications for production environments across diverse infrastructures.
Abstract
While major cloud providers and hyperscalers leverage virtually unlimited computational resources, the majority of organizations face significant constraints when operationalizing Large Language Models. Abdel Sghiouar presents a comprehensive set of practical strategies for optimizing LLM inference workloads on Kubernetes. The session systematically addresses container and model optimization techniques, accelerator management, data persistence and storage considerations, networking and intelligent load balancing, and advanced observability practices. Emphasis is placed on open-source tools and architectural patterns that deliver meaningful cost-performance improvements adaptable to on-premises, hybrid, and public cloud deployments.
Understanding LLM Inference Characteristics and Challenges
Large Language Models continue their rapid evolution in both scale and sophistication. Architectural innovations such as mixture-of-experts (MoE) enable dynamic activation of specialized sub-networks, while multi-modal capabilities process diverse inputs including text, images, audio, and video. Expanded context windows support richer interactions but demand substantial memory resources.
Inference execution comprises two primary phases with contrasting characteristics: the prefill stage (encoding input tokens, predominantly compute-bound) and the decode stage (token generation, typically memory-bound). KV (key-value) caching optimizes conversational flows by preserving intermediate states, avoiding redundant prefill computations for subsequent messages.
Deployment topologies vary considerably. Single-host single-accelerator setups predominate for local development and experimentation (e.g., using Ollama). Single-host multi-accelerator configurations require model sharding across GPUs within one machine. Multi-host distributed deployments introduce complex requirements for high-bandwidth, low-latency interconnects to maintain coherent context across nodes. Each topology presents distinct challenges regarding scalability, fault tolerance, and operational complexity.
Container, Model, and Storage Optimizations
Inference serving runtimes and model artifacts generate exceptionally large container images, frequently exceeding several gigabytes prior to incorporating weights. Conventional optimization strategies like multi-stage builds or native compilation (e.g., GraalVM) prove inadequate for these workloads.
Distributed caching solutions such as Spiegel provide cluster-wide image and model artifact caching, substantially reducing repeated pulls from external registries. Kubernetes-native features enabling containers as volumes allow separate packaging of models, which can then be mounted efficiently onto serving runtimes. When combined with caching layers, these approaches dramatically accelerate cold starts.
Quantization techniques offer another lever, reducing numerical precision (e.g., FP16 to INT8 or lower) to decrease memory footprints while preserving sufficient accuracy for many applications. Careful selection of quantization levels based on task sensitivity balances performance and quality.
Accelerator Management and Dynamic Resource Allocation
Kubernetes has supported GPU scheduling through device plugins for several years. However, static device configurations struggle with real-world constraints including accelerator scarcity and heterogeneous hardware fleets.
Dynamic Resource Allocation, matured in recent Kubernetes versions, introduces flexible resource claiming based on abstract characteristics rather than rigid device specifications (e.g., requesting “NVIDIA GPU with minimum 30GB memory and specific core count”). This enables more efficient scheduling across mixed clusters and better utilization rates.
Integration with cluster autoscalers allows on-demand provisioning, addressing both availability gaps and cost optimization by scaling resources precisely to workload demands. Platform operators describe device inventories; application teams specify requirements, with the scheduler performing intelligent matching.
Networking, Load Balancing, and Observability Considerations
LLM traffic profiles differ markedly from conventional web workloads. Requests exhibit high variability in size and computational intensity (simple text queries versus multi-modal inputs), while responses frequently involve streaming token generation. Standard round-robin load balancing produces inefficient distributions, with certain backends becoming overloaded while others remain underutilized.
The Kubernetes Gateway API, augmented with custom endpoint selection logic, supports sophisticated routing decisions based on request attributes extracted from bodies (model identifier, input modality, streaming requirements) combined with real-time backend telemetry. This facilitates intelligent traffic steering, prioritization of business-critical workloads, and maintenance of sticky sessions necessary for coherent streaming interactions.
Comprehensive observability must encompass prefill and decode phase latencies, KV cache hit rates, token generation throughput, GPU utilization, and end-to-end request metrics. Integration with Prometheus, Grafana, and specialized LLM monitoring solutions provides actionable insights for capacity planning and bottleneck identification.
Practical Patterns and the LLM-D Project
The LLM-D initiative, hosted under the Linux Foundation with contributions from Google, IBM, NVIDIA, and additional partners, aggregates architectural patterns, performance benchmarks, and reference implementations for production-grade inference. Key elements include optimized prefill/decode separation, advanced routing logic often leveraging engines like vLLM, and comprehensive guidance for multi-node deployments.
A holistic, layered optimization strategy proves most effective: infrastructure-level improvements (caching, persistent volumes), platform capabilities (dynamic scheduling, intelligent networking), and application-level choices (model quantization, serving engine selection). Organizations without hyperscale resources can still achieve competitive efficiency and scalability through disciplined application of these patterns.
Links:
[PyDataGlobal2025] Using Traditional AI and Large Language Models to Automate Complex and Critical Documents in Healthcare
Lecturer
Lily Xu is a Data Science Director in the corporate data-science and AI team at Vertex Pharmaceuticals, where she has worked for approximately seven years. She leads interdisciplinary groups of data scientists, data engineers, software engineers, and operations specialists focused on clinical-area solutions. She holds a doctorate in bioengineering from the Massachusetts Institute of Technology and an undergraduate degree from the University of California, Berkeley. Her earlier research produced publications on virtual microfluidics and the human microbiome; at Vertex she has driven projects spanning generative AI for clinical documentation, predictive patient modeling, large-scale claims analytics, protocol design, and centralized site intelligence.
Abstract
Informed consent forms constitute high-stakes, patient-facing, heavily regulated documents that must be tailored to jurisdictional requirements, local ethics boards, and plain-language standards. Their manual production across dozens of countries and hundreds of sites creates substantial operational bottlenecks in clinical-trial start-up. This article examines a production system developed at Vertex Pharmaceuticals that combines classical document-processing pipelines with large language models to auto-draft informed consent forms at scale. Emphasis is placed on architectural choices that minimize hallucination risk, rigorous measurement of end-to-end time savings, the centrality of change management, and the longer-term strategy of constructing a connected document network rather than isolated point solutions.
Clinical-Trial Operations Context and the Dual AI Portfolio
Clinical-trial operations span design, planning, execution, and monitoring phases, each generating or consuming large volumes of structured and unstructured documents. Failure to recruit patients, suboptimal site selection, or protracted regulatory review can each cost tens to hundreds of millions of dollars. Beginning in 2019 the Vertex data-strategy and solutions team—functioning as an internal SWAT unit—built trust through small, measurable pilots that combined public and private data into AI-ready assets. Over successive years the portfolio matured from ad-hoc analytics into standardized offerings for site identification, patient finding, enrollment forecasting, and, more recently, generative document automation.
The team deliberately distinguishes analytical AI (predictive modeling, Bayesian enrollment forecasts, rare-disease patient identification) from generative AI (first-draft document creation, knowledge-base chat, brand-copy generation). Business partners often approach the group believing a problem requires generative technology when structured data and classical machine learning would suffice; conversely, generative methods unlock previously intractable free-text tasks. Framing the two categories helps both data scientists and operational stakeholders select the appropriate tool. A foundational data layer aggregates site performance metrics, physician databases, claims, and census information; disease-specific analytic views and predictive models sit atop this foundation. Parallel generative pipelines extract structured content from lengthy protocols and feed downstream document generators, with embedded quality-control workflows so that extraction errors are corrected before they propagate into patient-facing material.
Architecture of the Informed-Consent-Form Auto-Drafting System
An informed consent form must convey risks, procedures, and rights in plain language while satisfying country-specific and sometimes site-specific regulatory requirements. A single multi-country trial may therefore require dozens of distinct variants. The solution developed at Vertex treats the clinical protocol as the primary source of truth, a blank regulatory template as the structural skeleton, and an approved language library as the repository of standardized phrasing.
Custom Python modules parse the protocol into logically coherent sections rather than arbitrary token chunks. Section-specific prompts and deterministic extraction routines pull the necessary facts. User-supplied answers to questions that cannot be parsed from the protocol are collected through a controlled interface. The resulting structured payload is inserted into the template; approved language snippets are retrieved via API from a purpose-built library that replaced earlier Excel spreadsheets and now maintains full audit trails and disease-area tagging.
The application is implemented in Flask and Dash, hosted on AWS behind single-sign-on, and calls a private Microsoft OpenAI endpoint for the generative steps. A monitoring dashboard continuously compares newly generated drafts against ground-truth forms produced by human experts, allowing the team to detect drift in accuracy over time. Because the generative component constitutes only a minority of the code base, the majority of engineering effort is devoted to robust parsing, template management, and workflow orchestration—skills that remain essential even as language models improve.
The design philosophy is “AI in the human loop” rather than “human in the AI loop.” Regulatory and patient-safety constraints demand that every draft undergo expert review; the system’s value lies in accelerating the initial drafting phase so that reviewers begin from a high-quality baseline rather than a blank page.
Measuring Impact, Change Management, and Scaling Strategy
Early controlled experiments compared pure manual drafting (one to three hours depending on trial complexity) with auto-draft generation (under ten minutes). Drafting-time reduction approached 90 percent. When subsequent editing and quality-control effort was included, net end-to-end time savings settled near 40 percent—still substantial given the volume of forms required across a growing portfolio. Because operational teams are chronically time-constrained, such measurements were performed on only two trials; the results nevertheless provided the quantitative foundation for continued investment.
Technology alone does not guarantee adoption. Change-management activities therefore received equal attention: standardization of templates and language libraries, transparent communication of model assumptions and known failure modes, and staged training that enabled business users to generate drafts independently. Treating free-text language assets with the same governance rigor applied to numerical data proved essential.
The longer-term vision is a connected document network rather than a collection of isolated point solutions. Clinical protocols and clinical study reports function as central hubs; mapping the full input–output relationships among start-up documents reveals opportunities for shared extraction components and cascading automation. The same platform is being extended to site budgets, case-report-form specifications, training materials, and other protocol-derived artifacts. Country-level templates are already linked so that a single protocol upload can spawn multiple jurisdiction-specific drafts simultaneously. Site-level customization remains outside the automated scope because the return on investment diminishes rapidly at that granularity; country-level guidance is instead provided to local teams.
Broader Lessons for Generative Applications in Regulated Environments
Several observations travel beyond the specific use case. First, impact measurement must be designed from the outset; without side-by-side timing studies and accuracy tracking, claims of productivity gain remain anecdotal. Second, the majority of engineering effort in production document systems continues to reside in classical software and data-engineering practices; large language models occupy a focused niche once reliable extraction and templating are in place. Third, alignment with business ownership is decisive: projects lacking motivated operational sponsors are deferred in favor of those with clear accountability and enthusiasm. Finally, the cumulative benefit of a systematically constructed document network can outweigh the initial development cost provided the organization persists past the early pilots.
Ambient listening, internal retrieval-augmented generation over institutional knowledge bases, and protocol optimization via real-world data are complementary initiatives already underway at Vertex and peer organizations. Collectively they illustrate a measured trajectory in which generative and analytical methods remove routine cognitive load while leaving critical reasoning and final accountability with domain experts.
Links:
[MiamiJUG] Retrieval-Augmented Generation: Building Deterministic AI for Production
Lecturer
Frank Greco is a Java Champion, enterprise architect, and senior consultant specializing in Artificial Intelligence and Cloud computing. He is the founder and Chairman of NYJavaSIG and a co-author of JSR #381 “VisRec,” the Java API for visual recognition. Frank is a recognized educator and technical leader who has presented at major global conferences including JavaOne, DevNexus, and Devoxx.
Abstract
This article provides an analytical framework for integrating Large Language Models (LLMs) into production Java environments using Retrieval-Augmented Generation (RAG). By moving beyond simple chat interfaces to programmatic API access, developers can build AI systems that are grounded in verified enterprise data. The analysis explores prompt engineering methodologies—such as Few-Shot and Chain of Thought (CoT)—and the architectural role of vector databases in mitigating model hallucinations while ensuring data security and version control.
Methodologies in Prompt Engineering
Prompting is the primary mechanism for steering the behavior of a neural network. Unlike traditional programming, prompting is probabilistic rather than deterministic. Frank identifies several advanced techniques to improve model reliability:
- Zero-Shot and Few-Shot Learning: Few-shot prompting provides the model with specific examples of the desired input-output pattern, significantly improving the accuracy of complex tasks.
- Chain of Thought (CoT): This instructs the model to “think step-by-step,” detailing its reasoning process before providing a final answer. This methodology is critical for reducing logical errors.
- Persona Identification: Assigning a specific role to the model (e.g., “Act as a Java security expert”) helps contextualize the response and refine the output tone.
Architectural Implementation: Retrieval-Augmented Generation (RAG)
To overcome the limitations of an LLM’s static training data, enterprises utilize RAG to ground the model in real-time, private data. In a RAG architecture, a user query is first used to search a knowledge base—typically a Vector Database—for relevant documents. This retrieved context is then injected into the prompt, allowing the LLM to generate an answer based on specific facts rather than general probabilities.
This approach offers several production-grade benefits:
- Reduced Hallucinations: By providing the model with the necessary facts, the likelihood of it “making up” information is significantly decreased.
- Data Security: RAG allows models to use private company information without that data being used to train the underlying public model.
- Traceability: Responses can be cited back to specific source documents found in the vector database.
Production Challenges and Ethical Considerations
Implementing AI at scale introduces significant engineering overhead. Developers must manage Prompt Versioning to ensure consistent behavior across deployments and navigate the legal implications of AI-generated content. Furthermore, because these are probabilistic systems, Frank warns that if a wrong answer poses a high risk to the business, generative AI may not be the appropriate solution. Engineers must balance the productivity gains of AI with the need for rigorous safety guardrails and human-in-the-loop verification.
Links:
[reClojure2025] LLMs + Clojure = Who needs frameworks?
Lecturer
Kapil Reddy is a software engineer known for his “business-first” approach to development. He is a prominent figure in the Clojure community, frequently contributing to discussions and ideation at the Scicloj meetups. Kapil has collaborated with other leading engineers in the ecosystem, such as Vedang Manerikar and Daniel Slutzky, to explore the intersection of artificial intelligence and functional programming. He is currently involved in developing the llms.edn project, which aims to bridge the gap between Clojure’s library-centric philosophy and the modern need for rapid project scaffolding using Large Language Models (LLMs).
Abstract
In the modern software development landscape, Large Language Models (LLMs) have significantly altered workflows, particularly in the realm of project scaffolding. However, the Clojure ecosystem, which prioritizes a philosophy of composable libraries over rigid frameworks, often presents a steep learning curve for newcomers who seek the convenience of “Rails-like” frameworks. This article explores a novel methodology introduced by Kapil Reddy that leverages LLMs to automate the composition of Clojure libraries. By utilizing a structured, native format called llms.edn, developers can describe library usage patterns in a way that LLMs can understand and execute. This approach aims to provide the convenience of a framework while maintaining the flexibility and power of Clojure’s traditional library-based architecture.
The Framework Paradox in Clojure
The debate between using frameworks versus a collection of libraries is central to Clojure’s identity. Traditional frameworks like Ruby on Rails provide a “Golden Path,” offering a set of pre-configured tools and conventions that allow for rapid prototyping. For many developers, especially those transitioning from other ecosystems, the absence of such a framework in Clojure is perceived as a significant barrier to entry. Clojure’s core philosophy leans heavily toward composition, where developers select specialized libraries—such as Ring for HTTP, Reitit for routing, and HugSQL for database access—and manually integrate them.
While this library-centric approach prevents the “black box” complexity and “magic” often associated with frameworks, it requires a deep understanding of the ecosystem. Kapil Reddy observes that LLMs are exceptionally proficient at project scaffolding, a task traditionally reserved for frameworks. The challenge, therefore, is to create a system where LLMs can assist in this scaffolding process without forcing the community to adopt a monolithic framework that would sacrifice the language’s fundamental strengths.
llms.edn: Structured Knowledge for AI Agents
To enable LLMs to effectively compose Clojure libraries, Kapil proposes a structured, Clojure-native approach to describing libraries and their common usage patterns: llms.edn. This concept is inspired by the broader llms.txt initiative but is tailored specifically for the unique requirements of the Clojure ecosystem.
The llms.edn file serves as a manifest that provides the LLM with the necessary context to understand how a library should be initialized, configured, and integrated with others. Instead of the LLM relying on potentially outdated or hallucinatory training data, llms.edn provides a source of truth directly from the library authors or the community. This structured data includes:
* Dependency declarations: Specific coordinates for tools like deps.edn or Leiningen.
* Code snippets: Standard boilerplate for starting a server or connecting to a database.
* Interoperability rules: Instructions on how a library (e.g., a router) interacts with another (e.g., a handler).
By providing these instructions in a machine-readable format, the manual task of “wiring” libraries together—often the most frustrating part for beginners—can be offloaded to an AI agent.
LLM-Powered Composition Workflows
The practical application of this methodology is an LLM-powered composition workflow. In this model, the developer describes the desired features of their application in natural language. An AI agent then queries a registry of llms.edn files to identify the best libraries for the task.
Kapil demonstrates that once the “how-to” for each library is codified, the process of generating a cohesive starter project becomes a “looper making a REST call”. This flow engineering treats the LLM as a pipeline that manages state and passes configuration data between different execution steps. This results in a “framework-like” experience where a full project structure is generated instantly, yet the underlying code remains a collection of simple, independent libraries that the developer can easily modify or replace.
The implications of this shift are profound. It suggests that the primary utility of a framework—reducing the cognitive load of setup and configuration—can now be achieved through intelligent automation. As Kapil notes, the LLM world requires more “simple software” because the models themselves introduce enough complexity; Clojure’s inherent simplicity makes it an ideal target for this kind of AI-driven orchestration.
Links:
[PyDataGlobal2025] What’s Next in AI for Data and Data Management
Lecturer
Lisa Amini is a Distinguished Engineer at IBM and Director of Data & AI Platforms Research, where she also leads IBM’s AI Horizons Network. Her career at IBM Research spans more than two decades and includes foundational work on stream processing systems that became the InfoSphere Streams product, leadership of the IBM Research laboratory in Ireland, and earlier roles directing knowledge and reasoning research. She has guided interdisciplinary efforts across cloud computing, artificial intelligence, and quantum computing, always with an emphasis on technologies that can be deployed at enterprise scale.
Abstract
Recent advances in large language models have catalyzed a wave of AI-assisted tools for data management and operations, ranging from code-generation assistants for data-flow pipelines to retrieval-augmented generation systems and increasingly autonomous data agents. This keynote examines the rapid evolution of generative and agentic capabilities, situates them within the broader data-management stack, and explores both near-term practical applications and longer-horizon research directions. Particular attention is given to the shift from human-operated systems augmented by copilots toward semi-autonomous stacks in which agents design, optimize, remediate, and continuously evaluate data products. The discussion balances technical opportunity with the enduring requirements of price-performance, open-source interoperability, and hybrid data architectures.
The Accelerating Capability Curve and the Emergence of Agency
The pace at which machine-learning benchmarks reach human-level performance has changed dramatically. Tasks that once required decades of incremental progress—handwriting recognition, for example—now reach parity within a few years or even months. Reading comprehension and predictive reasoning benchmarks follow similarly steep trajectories. While these evaluations remain narrow and do not constitute artificial general intelligence, they illustrate an unprecedented rate of improvement. Simultaneously, the cost per inference continues to fall even as model size and training compute grow, a trend driven by better systems design and algorithmic efficiency.
Within this landscape the progression from predictive models to generative models to conversational systems and finally to agents marks a qualitative shift. Agents do not merely answer questions; they dynamically control application flow, make decisions, take actions, and attempt self-correction. In the data domain this agency opens the possibility of systems that no longer wait for humans to formulate every query or repair every broken pipeline. Instead, agents can probe schema, resolve ambiguity, hypothesize data products, evaluate their own output, and iterate.
Transforming the Data Landscape and the Complementary Task Stack
Unstructured data has long existed, yet only recently has it assumed central importance. Machines can now reason over images, generate multimodal content, and extract structured signals from free text at scale. Classical database, warehouse, and lakehouse architectures, optimized primarily for structured tables, must therefore accommodate new access patterns. Retrieval-augmented generation pipelines replace static queries with dynamic retrieval-plus-generation cycles. User interaction moves from fixed application-generated SQL toward speculative, multi-step agent dialogues that probe metadata, formulate candidate queries, and refine them in light of intermediate results.
A useful conceptual inversion is to view the traditional storage–compute–query stack alongside a complementary human-task stack: infrastructure design, workload optimization, data discovery, enrichment, flow creation, remediation, governance, and insight generation. Each of these human activities constitutes fertile ground for agentic automation. Early systems already demonstrate learnable components inside query optimizers and routers; more ambitious research explores whether agents can search the design space of kernel-level software itself.
From Automation to Autonomy: Data Products and Continuous Evaluation
The practical goal is not merely to accelerate individual steps but to move entire workflows from human-operated to human-supervised. Consider the request to stand up a data stack and associated data products for a new application—robo-trading, for instance, that must combine public market data with sentiment signals and support periodic rebalancing. A multi-agent system can be tasked with discovering relevant sources, hypothesizing an ideal schema, populating that schema from heterogeneous tables and documents, extracting structured fields from natural-language text, and packaging the result as a governed data product.
Critical to autonomy is the ability to evaluate quality without constant human intervention. One effective strategy generates natural-language questions that a domain expert would plausibly ask of the intended data product, translates those questions into executable queries, and then monitors coverage metrics (tables and columns touched), topic coverage, query complexity, and latency. Agents iterate—adding sources, refining transformations, simplifying views—until the metrics stabilize within acceptable bounds or progress plateaus and human guidance is required. The same loop can later serve as continuous monitoring: questions that once succeeded can be re-executed to detect drift or regression.
Similar patterns apply to operational remediation. When a data-flow pipeline fails, agents can examine logs, generate natural-language root-cause hypotheses, propose script repairs, and, under appropriate guardrails, test those repairs in a sandbox before presenting them for approval. Across the spectrum of use, build, and optimize activities, the user’s role gradually shifts from operator to approver or observer.
Enduring Constraints and the Research Horizon
Price-performance remains non-negotiable; open-source components continue to enable rapid composition of storage formats, query engines, and table formats; hybrid architectures that span on-premises, cloud, and edge locations persist. Benchmarks, data contracts, open lineage standards, and carefully scoped open-weight models supply the interfaces and evaluation harnesses that allow agents to interoperate safely. Research prototypes already explore operator libraries that let developers request high-level transformations while large language models synthesize the concrete implementations behind the scenes.
The path forward is incremental. Fully autonomous data stacks will not appear overnight. Yet the combination of generative models, agent frameworks, and rigorous evaluation loops is already moving concrete workloads—data-product curation, flow repair, insight generation—along the continuum from assistance toward autonomy. The opportunity for data scientists and engineers is to shape the metrics, tools, and governance practices that will keep these systems both powerful and trustworthy.
Links:
[VoxxedDaysBucharest2026] Building a Sarcastic, Agentic Pair Programmer: Alexander Chatzizacharias on Crafting Playful LLM Workflows
Lecturer
Alexander Chatzizacharias is a software engineer at JDriven, a specialized consultancy in the Netherlands focused on JVM technologies and modern software development practices. With a unique background blending Dutch and Greek influences and a keen interest in game studies, Alexander brings creativity and playful thinking to technical challenges. He frequently speaks on topics including Java, Spring Boot, AI applications, and innovative development workflows.
Abstract
As mainstream AI coding assistants converge toward similar polished but somewhat generic experiences, Alexander Chatzizacharias demonstrates how to build a highly personalized, characterful AI pair programmer named “Pip.” Inspired by interactions with a sarcastic colleague named Ricardo, Pip incorporates personality through vectorized Slack history, utilizes Spring Boot and Kotlin, runs entirely locally with Qwen models via Ollama, and employs sophisticated workflows, multi-vector RAG, and the Model Context Protocol (MCP) to create delightful and productive assistance while addressing challenges like non-determinism and model drift.
The Homogenization of AI Assistants and the Quest for Personality
Alexander observes that leading AI coding tools have converged on remarkably similar chat-based interfaces and interaction patterns, largely influenced by OpenAI’s design choices. While incremental improvements continue, the overall experience feels increasingly uniform. This observation inspired the creation of Pip — an intentionally quirky, sarcastic AI pair programmer that injects personality drawn from real colleague interactions.
By processing Slack conversation history into vector embeddings stored in Qdrant, Pip can retrieve and emulate Ricardo’s characteristic sarcastic tone, witty retorts, and playful threats (such as threatening to delete poorly written code). This transforms the assistant from a neutral tool into a more engaging, human-like collaborator that questions unclear requirements, offers humorous feedback, and makes the development process more enjoyable.
Technical Architecture: Workflows, Agents, and Local Execution
Pip is implemented as a Spring Boot application written in Kotlin, with an IntelliJ IDEA plugin providing the frontend interface. Everything runs locally to maintain privacy and control: Qwen 3.5 models served through Ollama handle the language tasks.
Rather than pursuing fully autonomous agents, Alexander favors structured workflows that provide greater determinism and reliability — attributes particularly valued in enterprise environments. A categorization agent, functioning as an LLM-as-Judge, routes incoming queries to appropriate specialized handlers. Each handler uses carefully crafted system prompts derived from Slack history to consistently embody the desired personality traits.
The architecture incorporates multiple specialized agents for response generation, sophisticated RAG pipelines leveraging both dense and sparse vector representations with ColBERT reranking for improved retrieval quality, and integration with the Model Context Protocol (MCP) for tool usage such as playing music or generating memes when appropriate.
RAG, Tools, and the Challenges of Non-Determinism
Retrieval-Augmented Generation forms a cornerstone of Pip’s capabilities, dynamically pulling relevant context to overcome the inherent token limitations of even advanced models. Multi-vector search strategies combine semantic understanding with keyword precision for more reliable information retrieval from project documentation, codebases, and conversation history.
Tool integration via MCP enables rich interactions but introduces additional complexity due to the non-deterministic nature of local models. Alexander discusses practical challenges including prompt sensitivity to model updates (“model locking” strategies), the art of prompt engineering which he likens to “vibe checking,” and the necessity of implementing guardrails to maintain appropriate behavior boundaries.
Implications for Future AI Development
Alexander encourages attendees to experiment with building personalized, domain-specific AI assistants using accessible open-source tools. While acknowledging the increasing commercialization of AI, he emphasizes the current window of opportunity for creative, playful implementations that enhance both productivity and developer satisfaction.
Pip serves as an inspiring example of how thoughtful combination of RAG techniques, vector databases, workflow orchestration, and personality injection can create AI tools that feel genuinely collaborative rather than merely functional.
Links:
[reClojure2025] Writing Model Context Protocol (MCP) Servers in Clojure
Lecturer
Vedang Manerikar is the founder of Unravel.tech and a veteran software architect with over 15 years of experience in the Clojure ecosystem. Previously serving as the Head of Backend Engineering at Helpshift, Vedang has managed large-scale distributed systems and led complex technical migrations. At Unravel.tech, his work focuses on the intersection of Clojure and Artificial Intelligence, specifically building “Agentic Systems” and implementing Generative AI (GenAI) and Large Language Model (LLM) solutions. He is the author of mcp-cljc-sdk, a cross-platform Clojure SDK for the Model Context Protocol.
Abstract
The rapid advancement of Artificial Intelligence has created a need for standardized communication between AI agents and external systems. The Model Context Protocol (MCP), introduced by Anthropic, has emerged as a solution to the integration problem, providing a common interface for agents to interact with diverse data sources and tools. This article explores the architecture of MCP and argues that Clojure is uniquely positioned as an ideal language for implementing MCP servers. We analyze the protocol’s similarity to the Language Server Protocol (LSP), examine real-world applications in browser automation and communication platforms, and discuss how Clojure’s REPL-driven development and data-centric philosophy streamline the creation of powerful, composable AI workflows.
The Model Context Protocol: A New Standard for AI UX
At its core, MCP is an open standard designed to enable AI applications—such as Claude Desktop or Cursor—to access the external world in a structured manner. While one might ask why standard HTTP interfaces are insufficient, the answer lies in the integration problem. Without a standard, every AI agent would need a custom integration for every service (PostgreSQL, Google Drive, GitHub, etc.). MCP solves this by acting as a “USB port” for AI; developers write a server for their service once, and it becomes immediately accessible to any MCP-compliant agent.
Vedang describes MCP not just as a data access layer, but as a “baseline AI UX.” It defines how an agent discovers tools, reads resources, and follows prompts. This standardization allows for the creation of sophisticated workflows where an agent can, for example, use a Playwright MCP server to browse Hacker News, a WhatsApp MCP server to read messages, and a local filesystem server to summarize information and save it to a document. By providing a consistent interface, MCP shifts the focus from integration plumbing to the design of the agent’s behavior and user experience.
Clojure as the Premier Language for MCP
Clojure’s technical characteristics align remarkably well with the requirements of building MCP servers. The protocol is heavily reliant on JSON-RPC and the exchange of structured data, which plays directly into Clojure’s “data-as-code” philosophy. Vedang highlights several key reasons why Clojure developers are particularly well-prepared for the LLM world:
1. REPL-Driven Development: MCP servers often act as intermediaries between non-deterministic LLMs and deterministic systems. The ability to interactively test and refine server responses in a live REPL mirrors the iterative nature of working with AI.
2. Data Transformation: Clojure’s rich library for manipulating maps and vectors makes it trivial to transform complex API responses into the simplified “Context” required by LLMs.
3. Cross-Platform Capability: With the mcp-cljc-sdk, developers can write server logic once and deploy it on both the JVM (using clojure.main or GraalVM native images) and Node.js (via ClojureScript), providing flexibility in how the server is hosted and consumed.
Code Sample: Defining a Simple MCP Tool
(defmethod handle-request "tools/call" [request]
(let [{:keys [name arguments]} (:params request)]
(case name
"get-weather" (let [city (:city arguments)]
{:content [{:type "text"
:text (str "The weather in " city " is sunny.")}]})
{:error "Tool not found"})))
Practical Applications and Agentic Workflows
The power of MCP is best demonstrated through real-world “Agentic” use cases. Vedang shares examples of servers he has developed to automate complex tasks. One such server integrates with WhatsApp, allowing an AI agent to scan chat groups for business leads. Instead of a human manually reading hundreds of messages, the agent uses the MCP server to fetch the latest messages, identifies intent, and provides a summarized report of actionable items.
Another significant application is in browser automation. Using an MCP server for Playwright, an AI can navigate the web as a user would—logging into sites, extracting data from dynamically rendered pages, and performing actions. This allows for prompts like “Find me a hotel within walking distance of the reClojure conference,” where the agent autonomously searches maps, checks availability, and compares prices. These examples illustrate how MCP enables the transition from simple chatbots to true “agents” capable of multi-step reasoning and interaction with the physical or digital world.
The Future of Content-Centric AI
Looking ahead, the evolution of MCP suggests a shift toward a “content-is-king” paradigm. Current AI interactions are often limited by the UX of the chat box. However, with MCP, the focus can move toward the actual content being produced or modified—whether that is a codebase, a spreadsheet, or a document. Vedang envisions a future where multiple coding agents can work in parallel on the same repository, coordinated through a “better Git” or similar bidi-rectinal communication protocols enabled by MCP.
By standardizing the way agents interact with our tools, MCP paves the way for a new generation of software that is designed from the ground up to be AI-enhanced. For the Clojure community, this represents a significant opportunity to lead the development of the “AI UX” by building robust, composable servers that unlock the full potential of Large Language Models.
Links:
[PyDataGlobal2025] Enhancing Apache NiFi 2.x with Python Processors
Lecturer
Timothy Spann is a Senior Solutions Engineer at Snowflake. He brings extensive experience in generative AI, large language models, Apache NiFi, Kafka, Pulsar, Flink, Spark, and related streaming and big-data technologies. Previously he held developer-advocate and field-engineering roles at Cloudera, StreamNative, Hortonworks, and other organizations. He maintains an active open-source presence and regularly publishes practical examples of NiFi processors.
Abstract
Apache NiFi provides a visual, highly configurable environment for building data-flow pipelines. Version 2.x introduces first-class support for Python processors, allowing developers to embed arbitrary Python logic—including rich libraries for machine learning, natural-language processing, and geospatial conversion—directly into streaming workflows. This article describes the architecture of Python processors, the packaging and deployment process, representative use cases ranging from image captioning to real-time transit-data conversion, and the operational advantages of running such processors inside a managed NiFi environment such as Snowflake Openflow.
NiFi Fundamentals and the Value of Python Integration
NiFi is a visual tool that lets users drag, drop, and connect processors to form directed data flows. It natively handles hundreds of sources and sinks, maintains detailed lineage and audit trails, and offers flexible error-handling and back-pressure mechanisms. Data are stored in content and attribute repositories that support interactive inspection and replay. Because NiFi already excels at integration, the addition of Python processors removes the need to re-implement sophisticated logic in Java or to off-load processing to external Spark or Flink clusters for many enrichment tasks.
A Python processor follows a simple contract. The developer supplies a class that declares dependencies, performs optional initialization, and implements a transform method. The method receives a flow-file (the unit of data moving through the pipeline) together with its attributes, may inspect or modify content, may add or alter attributes, and returns the flow-file for downstream routing. Packaging produces a NAR archive that is dropped into a NiFi extension directory or uploaded through a managed interface such as Openflow. Once loaded, the processor appears in the palette exactly like any native component and can be parameterized, scheduled, and monitored through the ordinary NiFi user interface.
Representative Processors and Demonstration Workflows
Concrete examples illustrate the range of possibilities. An RSS reader built on feedparser and pandas ingests government news feeds and emits CSV. An image-captioning processor loads a Hugging Face BLIP model, receives an image flow-file, writes a natural-language caption into an attribute, and passes the original image unchanged. Subsequent processors can apply ResNet-50 classification or NSFW detection without copying the binary content. Named-entity recognition with spaCy extracts organizations and persons from text; an OpenStreetMap geocoder converts postal addresses into latitude-longitude pairs; a GTFS-realtime converter transforms Protocol-Buffer transit feeds from the New York MTA into JSON.
In a live Openflow demonstration a GTFS processor is configured with a URL and a feed type (trip updates, vehicle positions, or alerts). Data flow through the processor, emerge as structured JSON, are split, attribute-extracted, merged, and finally loaded into Snowflake tables—all without leaving the NiFi canvas. Because processors can be started, stopped, and reconfigured while the flow remains active, developers obtain an interactive feedback loop that is difficult to replicate in batch-oriented environments.
Operational Considerations and Broader Implications
Python processors run as external processes outside the JVM; consequently they are subject to different resource constraints and are typically restricted to medium or large runtime sizes in managed offerings. Best practice therefore reserves them for tasks that genuinely benefit from the Python ecosystem—model inference, specialized parsing, or rapid prototyping—while leaving high-volume, CPU-bound work to native Java processors. Parameterization separates sensitive or environment-specific values from version-controlled flow definitions, facilitating promotion across development, test, and production clusters.
The combination of NiFi’s integration strengths with Python’s analytic libraries yields a pragmatic architecture for real-time enrichment pipelines. Unstructured data—images, archives, Protocol-Buffer streams—can be ingested, enriched with machine-learning metadata, and routed to downstream systems such as Kafka, Iceberg tables, or Slack channels. The same pattern supports prompt construction and calls to external large-language-model endpoints, positioning NiFi 2.x as a convenient orchestration layer for hybrid streaming and generative-AI workloads.