Recent Posts
Archives

Posts Tagged ‘Robotics’

PostHeaderIcon [AWSReInvent2025] Control Humanoid Robots and Drones with Voice and Agentic AI

Lecturer

Hang Celia is a developer advocate at Amazon Web Services (AWS) based in Hong Kong, specializing in AI and robotics integrations. Saras Wang is a senior AWS Hero from Hong Kong, actively contributing to social media platforms and community discussions on cloud technologies.

Abstract

This article investigates the integration of voice control with agentic AI for managing humanoid robots, robot dogs, and drones, drawing from a collaborative project with the Hong Kong Institute of Information Technology (HKIIT). It examines the architecture for low-latency command processing, intent recognition, and responsive behaviors, while analyzing methodologies for handling continuous speech and multi-robot coordination, along with their broader implications for real-world applications.

Overview of Agentic AI and Its Future Predictions

Agentic AI marks a significant advancement in the field of artificial intelligence, shifting from passive response systems to proactive entities capable of independent planning, decision-making, and execution of complex tasks in dynamic settings. Hang Celia sets the stage by drawing on insights from leading investment analyses, which project a profound impact on various industries. For example, Goldman Sachs anticipates that by 2027, agentic AI could automate as much as 25% of routine work activities, thereby reshaping labor markets and boosting productivity across sectors. Similarly, McKinsey’s projections suggest that by 2030, this technology might account for 30% of current work hours, highlighting its potential to revolutionize operational efficiencies, especially in areas demanding real-time adaptability such as automated systems and robotics.

Building on these forecasts, agentic AI extends beyond traditional large language models by incorporating advanced capabilities like logical reasoning, external tool integration, and iterative problem-solving over multiple stages. Hang illustrates this evolution through practical demonstrations, where an agent might receive a natural language command, break it down into actionable components, query external resources via APIs, and refine its approach based on ongoing feedback. This stands in stark contrast to earlier AI paradigms, which were largely reactive and limited to single-turn interactions, and instead positions agentic systems as versatile facilitators for sophisticated human-machine collaborations, particularly in controlling physical devices like robots.

The underlying methodology for deploying agentic AI in such contexts relies heavily on cloud-based services, with AWS offerings like Amazon Bedrock providing the orchestration layer that enables seamless access to knowledge repositories and function executions. This not only facilitates rapid prototyping but also ensures that the systems can scale to handle diverse inputs and outputs. Consequently, the implications are far-reaching, as agentic AI holds the promise of making advanced robotic controls more intuitive and widespread, extending their utility from specialized research environments to everyday applications in homes, offices, and industrial facilities.

Architecture for Voice-Controlled Robotics

The architectural design of the voice-controlled robotics system is engineered to support seamless and natural interactions, combining speech processing, natural language comprehension, and agentic execution to achieve responses with minimal delay and maximal accuracy. Saras Wang provides a detailed walkthrough of the system’s structure, which harnesses a suite of AWS services to transform spoken commands into precise directives for a variety of robots, including humanoids, quadruped models, and aerial drones. At its core, the setup begins with Amazon Transcribe, which converts audio streams into text in real time, enabling the system to interpret ongoing conversations without requiring artificial pauses or structured phrasing.

From there, the processed text feeds into Amazon Bedrock, where intent detection occurs, identifying the user’s objectives and mapping them to specific robot functions. This integration allows for flexible handling of commands, such as directing a humanoid to perform a gesture while simultaneously instructing a drone to adjust its position. Saras emphasizes the importance of WebSockets in maintaining bidirectional communication channels, which facilitate not only command issuance but also feedback loops from the robots, ensuring that the system can adapt to changing conditions or confirm task completions.

In terms of methodology, the approach prioritizes optimization for diverse environments, incorporating noise-reduction algorithms to filter out background interference and edge computing elements to minimize latency in transmission. Challenges like varying accents or ambiguous phrasing are addressed through machine learning models trained on extensive datasets, which refine recognition over time. Overall, this architecture enhances usability by making robotic control as intuitive as everyday speech, while its modular design supports expansions to new device types or additional functionalities without overhauling the core framework.

Multi-Robot Coordination and Parallel Execution

Coordinating actions across multiple robots introduces layers of complexity in terms of synchronization and resource allocation, yet the project demonstrates effective solutions through strategic function calling and API optimizations that enable simultaneous operations. Hang elaborates on how agentic AI can trigger parallel invocations, allowing a single voice command to engage several devices without sequential bottlenecks. For instance, a directive to have all robots rotate could be decomposed, with the agent assigning unique tasks to each unit—perhaps turning one left, another right, and a third forward—while ensuring no conflicts in shared spaces.

Saras offers practical code insights to illustrate this parallelism:

import concurrent.futures

def control_robot(robot_id, action):
    '''# API call to robot'''
    response = robot_api.execute(robot_id, action)
    return response

with concurrent.futures.ThreadPoolExecutor() as executor:
    future1 = executor.submit(control_robot, 'robot1', 'turn_left')
    future2 = executor.submit(control_robot, 'robot2', 'move_forward')
    results = [future1.result(), future2.result()]

This code leverages threading to execute commands concurrently, significantly reducing overall response times. The methodology involves designing robot APIs to support asynchronous calls, with AWS Lambda or similar services handling orchestration to distribute loads evenly. In real-world contexts, this prevents overloads during high-demand scenarios, such as coordinated search operations with drones and ground robots.

The implications for scalability are substantial, as this framework can extend to fleets of dozens or hundreds of units, applicable in logistics warehouses or disaster response teams. By prioritizing parallel processing, the system not only improves efficiency but also enhances reliability, as failures in one robot do not halt the entire operation.

Challenges, Innovations, and Real-World Implications

While the fusion of voice interfaces with agentic AI offers immense promise, it also surfaces obstacles like debugging intricate integrations and managing network dependencies, which the project overcomes through iterative innovations and tool leveraging. Saras reflects on initial hurdles: early attempts avoided frameworks for perceived simplicity, but this led to unresolved issues in error handling and scalability. Transitioning to structured frameworks, such as AWS CLI for API conversions, resolved these, underscoring the importance of utilizing pre-existing solutions to address common pitfalls without reinventing foundational elements.

Innovations include adapting request-response APIs to streaming formats for continuous dialogues, facilitated by Amazon Q’s automation capabilities. Hang notes experiments with digital humans, where APIs process multilingual documentation—such as simplified Chinese sources—via AI-driven implementations, broadening accessibility.

Broader real-world implications span from educational tools, where students command robots intuitively, to assistive technologies for the elderly, enhancing independence. Future enhancements might include office automation, where voice directives control devices seamlessly, transforming how humans interact with intelligent systems in daily life.

Conclusion

The HKIIT-AWS collaboration vividly demonstrates how agentic AI and voice control can elevate robotics to new levels of practicality and engagement. By tackling coordination challenges and harnessing AWS infrastructure, it establishes a foundation for innovative applications that bridge the gap between human intent and machine action.

Links:

  • https://www.youtube.com/watch?v=ZKqV1Ok-2-c

PostHeaderIcon [DevoxxGR2026] Code That Moves the World: The Rise of Physical AI

Lecturer
Will Sentance is the founder of Standard Material and Codesmith, organizations at the forefront of physical AI infrastructure and AI/software engineering education. A speaker, educator, and practitioner, Sentance bridges software engineering expertise with emerging robotics and autonomous systems. He contributes to research at Oxford and leads initiatives training talent for the next wave of intelligent physical systems.

Abstract
In this forward-looking keynote at Devoxx Greece 2026, Will Sentance explores the profound convergence of software engineering and physical intelligence. Robots and autonomous systems are transitioning from specialized, brittle demonstrations to capable, generalizable agents operating in real-world environments. Sentance details the technological breakthroughs in hardware, data, and foundation models driving this transformation and argues that traditional software engineering skills are central to building the platforms, data pipelines, and integrations required for scalable physical AI deployment.

The Remarkable Progress in Physical Intelligence

Physical AI—systems that sense, understand, and act upon the physical world—has advanced dramatically. Robots now follow natural language instructions, handle novel objects, and demonstrate emergent capabilities. Foundation models for robotics enable zero-shot generalization and long-horizon planning across diverse embodiments.

Companies like Physical Intelligence, Agility Robotics, and others are moving from laboratory experiments to industrial and domestic applications. This shift is fueled by massive investment and rapid iteration.

Core Technological Enablers

Three key areas have transformed the landscape:

Hardware Revolution: Affordable, off-the-shelf components—from full humanoids to grippers and sensors—dramatically lower barriers. Edge computing platforms provide sufficient power for onboard inference.

Data Explosion: Teleoperation, simulation (including sophisticated world models), and real-world deployment generate multimodal datasets at unprecedented scale. Techniques like action chunking address real-time requirements.

AI Models: End-to-end learning replaces traditional control theory. Vision-language-action models predict continuous action trajectories, enabling flexible behavior without exhaustive manual programming.

The Physical AI Technology Stack

Sentance outlines a layered architecture:

  • Real-time Control: Low-level, deterministic operations managing actuators and safety at high frequency.
  • Platform and Middleware: Abstractions like ROS providing integration, simulation interfaces, and developer tools.
  • Intelligence Layer: Foundation models processing vision, language, and proprioception to generate actions.
  • Data and Learning Loop: Continuous collection, training, evaluation, and deployment cycle.

Opportunities for Software Engineers

Contrary to initial impressions, software engineers are perfectly positioned to lead this revolution. Approximately 80% of the required work involves familiar disciplines: systems architecture, platform engineering, data pipelines, low-level optimization, and agentic integration.

Roles at leading organizations emphasize scalable frameworks, reliable deployment, observability, and integration of AI models into production—skills honed in cloud-native and distributed systems development.

New challenges center on real-time constraints, physical dynamics, and managing massive multimodal datasets, but these build directly upon existing expertise.

Getting Started with Physical AI

Sentance encourages practical experimentation using affordable hardware like the SO-101 and open tools. Developers can quickly train policies for simple tasks such as closing a laptop lid, experiencing the full cycle from data collection to deployment.

The physical world represents the next major platform for code. Software engineers who embrace this frontier will shape the coming industrial transformation.

Links:

PostHeaderIcon [DotJs2025] Code in the Physical World

The chasm between ethereal algorithms and tangible actuators has long tantalized technologists, yet bridging it demands more than simulation’s safety nets— it craves platforms that tame the tangible’s caprice. Joyce Lin, head of developer relations at Viam, bridged this divide at dotJS 2025, chronicling how open-source orchestration empowers coders to infuse IoT and robotics with JS’s fluidity. A trailblazer in hardware-software symphonies, Joyce demystified the real world’s rebellion against unit tests, spotlighting Viam’s registry as a conduit for browser-bound brains commanding distant drones.

Joyce’s epiphany echoed Rivian’s rueful recall: OTA firmware’s folly, bricking 3% of fleets via certificate snafus—simulation’s simulacrum shattered by deployment’s deluge. The physical’s peculiarities—unpredictable pings, sensor skews, mechanical murmurs—defy CI/CD’s certainties; failures fleck the field, from rover ruts to vacuum voids. Viam’s virtue: a modular mosaic, JS SDKs scripting behaviors atop a cloudless core. Joyce vivified with vignettes: a browser dashboard dispatching drone dances, logic lingering in tabs while peripherals pulse commands via WebSockets. Serial symphonies follow: laptop-launched loops querying quadrature encoders, fusing firmware’s fidelity with JS’s finesse.

This paradigm pivots potency: core cognition—path plotting, peril parsing—resides in reprovable realms, devices demoted to dutiful doers. Viam’s vista: modular motions, from gimbal glides to servo sweeps, orchestrated sans silos. AI’s infusion amplifies: computer vision’s vintage, now vivified by low-cost compute—models marshaled, fleets federated, data’s deluge distilled into adaptive arcs. NASA’s pre-planned probes pale beside this plasticity; vacuums’ vacuums evolve, shelves’ sentinels self-optimize.

Joyce’s jubilee: tech’s tangible thrust—from wearables’ whispers to autonomous autos—blurs bytes and brass. Viam’s vault: docs delving devices, SDKs summoning synths—inviting artisans to animate the ambient.

From Simulation to Sentience

Joyce juxtaposed Rivian’s reckoning with Viam’s resilience: OTA’s overreach underscoring physicality’s pitfalls—cert snares, signal storms. Browser-bound bastions: WebRTC webs weaving commands, logic liberated from latency’s lash.

Orchestrating the Observable

Viam’s vernacular: registries routing routines, JS junctions juggling joints—gimbal gazes, encoder echoes. AI’s ascent: models’ maturity, compute’s cascade—rover reflexes refined, vacuum vigils vivified.

Links: