Research
Physical AI Still Trapped Between Simulation and Reality, Robotics Developers Say
At the Actuate conference, robotics engineers described the field as stuck in a GPT-2-era stage, where abundant enthusiasm masks fundamental gaps in data quality, simulation fidelity, and production reliability. The constraint is not theory but engineering: companies can build robots that work in controlled settings, but scaling to commercial deployment remains blocked by the same bottlenecks that have limited autonomous systems for years.
By Michael G ·

The robotics industry is experiencing a contradiction familiar to anyone who watched the autonomous vehicle boom: impressive demonstrations paired with stubborn engineering limits. According to reporting from the TechCrunch report on the Actuate conference held in August 2026, developers increasingly use the GPT-2 comparison to describe where physical AI stands. The analogy is precise. GPT-2 was competent enough to generate coherent text but lacked the scale, refinement, and reliability of later models. Today's robots can navigate tasks in lab conditions and controlled warehouses, but they fail unpredictably in the messier, more varied conditions that define real commercial work. The Actuate conference drew roughly 1,500 attendees and has tripled in size since 2023, reflecting genuine industry momentum. Yet the size of the crowd masks the narrowness of deployed solutions.
The gap between capability and reliability hinges on three interconnected problems that developers identified repeatedly at the conference. First is data scarcity. Training modern robotic systems requires enormous volumes of task-specific, high-quality data capturing real-world variation. Unlike text or images, which exist in vast quantities online, robot training data must be either collected through costly real-world trials or synthesized through simulation. Most companies lack the resources or operational scale to generate sufficient data independently. Second is simulation quality. Physical robots interact with materials, friction, deformation, and environmental noise in ways that are extraordinarily difficult to model accurately. Simulators that work well enough for prototyping often fail at scale because small inaccuracies compound. Third is the feedback loop between deployment and improvement. Autonomous vehicles have partly overcome these constraints because companies like Uber operate massive human-driven fleets that generate relevant data continuously. Each trip produces signals about what went wrong and why, feeding back into the learning pipeline. Robotics companies lack equivalent data streams, making production deployment a catch-22: they need deployment data to improve, but cannot deploy reliably without better models. TechCrunch report provides the primary public record for that part of the account.
Autonomous vehicles illustrate the point. After more than a decade of development, companies like Wayve have demonstrated that vision-based driving can work in real cities under real conditions. Yet this progress came partly through infrastructure that robotics lacks. Human drivers generate constant feedback about edge cases, failures, and scenarios that trainers never anticipated. When a Wayve vehicle encounters an unexpected traffic pattern or road construction, the system learns from it. Robotics companies cannot replicate this at the necessary scale. A warehouse robot that fails on a task it has never seen before does not automatically contribute data for retraining; someone must debug, annotate, and incorporate the failure. The labor cost of this process remains prohibitively high for most applications. Foxglove supplies additional technical context for evaluating the claim.

Deployment Strategy and Commercial Reality
Given these constraints, the industry has bifurcated into two distinct strategies. Some companies, including Uber, have invested in humanoid robotics research as a longer-term bet on general-purpose manipulation. Others have focused on task-specific deployment in environments where the variation is manageable. Companies like Agility Robotics have moved beyond research into industrial deployment, placing robots in actual production settings where the task scope is narrow and controlled. This pragmatic approach trades generality for reliability. A robot designed specifically for warehouse pick-and-place or manufacturing assembly can achieve acceptable performance because the input space is constrained. The robot sees primarily the variation it was trained for. But this strategy scales slowly. Each new task requires expensive retraining and domain expertise. The path to economically viable robotics depends on solving the constraint problems rather than working around them. The operating constraint is also visible in material published by Nvidia Cosmos.
Technical infrastructure improvements are emerging, though not yet at scale. Foxglove announced a search product built on Nvidia Cosmos, representing an attempt to make better simulation tools accessible. Better simulation could reduce the data collection burden by generating more accurate synthetic training examples. Nvidia Cosmos aims to produce video prediction models that capture physical dynamics more faithfully than current approaches. If successful, this could accelerate the feedback loop between simulation and real-world testing. However, simulation quality improvements alone will not solve the problem. A better simulator still requires real-world validation and calibration. The fundamental constraint remains the same: moving from laboratory demonstration to reliable commercial operation requires both better tools and the labor to use them correctly.
The technical challenges extend beyond simulation and data into embodiment and control. Different robot designs have fundamentally different capabilities and failure modes. A humanoid robot moving through an office faces different constraints than a wheeled system designed for warehouses. The gripper design, joint configuration, and actuator speed all affect what tasks are practical and how quickly a system can be retrained. This embodiment problem is not new, but it becomes more acute as the industry moves from research toward deployment. Each design choice locks in assumptions about the tasks the robot will perform. Changing tasks requires not just new training data but potentially new hardware. Companies building task-specific systems accept this trade-off. Companies pursuing general-purpose humanoids must solve embodiment across a much wider range of scenarios, which is technically harder and requires substantially more training data. For institutional context, Wayve explains the relevant system or standard.

The MLOps and Deployment Economics Problem
Beyond the technical constraints lies an economics problem that the industry has not yet solved. Deploying and maintaining machine learning systems at scale requires mature MLOps infrastructure: continuous monitoring, automated retraining pipelines, careful version control, and robust testing. Software companies have built these systems over years. Robotics companies are mostly still improvising. A deployed robot that begins to fail cannot wait days for human engineers to investigate. The system must either fail gracefully or notify operators immediately. This requires instrumentation, logging, and alerting infrastructure that most robotics companies lack. When a robot fails in production, determining whether the failure stems from hardware wear, environmental change, model drift, or a novel scenario requires careful engineering. Without this infrastructure, deployment becomes risky and expensive. Companies must either maintain large on-site teams to monitor robots or accept downtime and lost productivity. Neither option scales well. Agility Robotics offers a separate reference point for the implementation question.
The evaluation problem compounds the deployment challenge. In autonomous driving, companies can measure success reasonably well: miles driven safely, disengagements per mile, incident rate. In robotics, defining success is harder because success depends on the specific task and the production environment. A manipulation robot's error rate in the lab may differ drastically from its error rate after six months in a factory. Without standardized evaluation frameworks, it is difficult to compare systems or to measure progress within a company over time. Research institutions like NIST robotics research have attempted to develop benchmarks, but these remain limited in scope and do not capture the complexity of real deployments. The industry lacks both the standardized evaluation methods and the data infrastructure needed to support reliable deployment at scale.
The feedback loop problem is ultimately what separates physical AI from reliable commercial work. In software, feedback is fast and cheap. An update reaches millions of users instantly. If it fails, revert and redeploy. In robotics, feedback is slow and expensive. A robot deployed in a factory operates in isolation. When it fails, the failure must be captured, communicated, and incorporated into retraining. This requires human effort, infrastructure, and time. The difference matters. A language model can improve from millions of user interactions per day. A deployed robot might provide feedback from dozens of interactions per day, assuming any feedback is captured at all. Closing this gap requires solving the data collection, simulation, embodiment, and deployment economics problems simultaneously. None of them has a simple solution. The enthusiasm at the Actuate conference reflects genuine technical progress. But progress in demonstrating capability is not the same as progress toward commercial reliability. The robotics industry remains in the stage where impressive demonstrations mask the grinding work of engineering reliability at scale. The unresolved issue can be assessed against guidance from NIST robotics research.
Topics: robotics, physical AI, research, manufacturing, autonomous systems