Research library · updated 2026-06-13 · public

Robotics robot-foundation-model data stack 2026 evidence artifact v1

Date: 2026-06-13 Owner: Finance / Charlie AGT-002 Visibility: PUBLIC Target: site + slide Status: source-backed artifact for Codex packaging Public-safety: no portfolio data, no trade recommendation, no private channel checks, no paid-report excerpts

0. One-line answer

The most valuable stale gap after the latest robotics research queue is the robot foundation-model data stack: Google DeepMind and NVIDIA are no longer only showing robot demos; they are exposing developer-facing layers for on-device VLA models, embodied reasoning APIs, open / downloadable humanoid foundation models, synthetic trajectory generation, teleoperation, simulation, and robot-specific post-training. This is a stronger S4 platform-formation signal than a single demo, but still not S5 commercial-economics proof because public sources do not disclose robot-model ARR, attach rates, gross margin, deployed fleet economics, uptime, intervention-rate improvement, or customer ROI/payback. 🟢 Google DeepMind / NVIDIA / Hugging Face primary sources; 🟠 Charlie stage classification.

1. Core question

If humanoid robotics scales, will the scarce layer be robot hardware, customer deployments, or the data/model stack that turns real-world experience into reusable capabilities?

Short answer as of 2026-06-13:

  • Hardware still matters, but the AI bottleneck is moving toward data generation, adaptation, embodied reasoning, and safe local inference. 🟢
  • The public model layer is becoming more legible because Google and NVIDIA now disclose specific developer-facing model/workflow primitives: Gemini Robotics On-Device, Gemini Robotics-ER 1.6, Isaac GR00T, GR00T-Dreams, GR00T N1.5, Isaac Teleop, Isaac Lab, Isaac Sim, Omniverse, Cosmos, and Jetson Thor. 🟢
  • This does not yet prove a standalone software-value-capture layer. The missing S5 evidence is paid usage, recurring revenue, contract economics, robot attach rate, gross margin, retention, customer ROI, or audited segment disclosure. 🟢 absence in reviewed sources / 🟠 inference.

2. Why this is additive

Existing robotics files already cover:

  • Tesla / Figure / Unitree / UBTECH / Agility / Apptronik / 1X company evidence curves.
  • BMW customer-side deployment evidence.
  • Logistics commercialization benchmarks.
  • Safety / standards gates.
  • Physical Intelligence and Skild AI as model-layer evidence.

This artifact adds the hyperscaler / platform-stack layer:

  • Google DeepMind: VLA + embodied-reasoning + on-device / low-latency developer path. 🟢
  • NVIDIA: open humanoid foundation model + simulation / synthetic data / teleoperation / deployment compute stack. 🟢
  • Public implication: robotics evidence should track not only robot counts and BOM, but also data-stack productivity: demonstrations required, synthetic-data generation time, model availability, embodiment coverage, local inference, and post-training workflow. 🟠 synthesis.

3. Evidence table

LayerSource-backed factQuantified / dated anchorWhat it changesSource gradeSignal grade
Google VLA on-deviceGoogle DeepMind introduced Gemini Robotics On-Device as a VLA model optimized to run locally on robots for low-latency inference and robustness where connectivity is intermittent or unavailable.Blog dated 2025-06-24.Moves the model question from cloud-only demos to edge deployment constraints: latency, connectivity, and local inference.🟢 Google DeepMindS4 platform signal
Google fine-tuning pathGoogle says Gemini Robotics On-Device is the first Gemini Robotics VLA model it is making available for fine-tuning; developers can adapt it with as few as 50-100 demonstrations.50-100 demonstrations; private preview / trusted testers.Gives a measurable adaptation-data anchor for robotics model evaluation.🟢 Google DeepMindS4 data-efficiency signal
Cross-embodiment adaptationGoogle says it trained the on-device model only for ALOHA robots but further adapted it to a bi-arm Franka FR3 and Apptronik Apollo humanoid.ALOHA -> Franka FR3 / Apptronik Apollo; 2025-06-24.Supports cross-embodiment ambition, but only as company-reported technical evidence.🟢 Google DeepMindS3/S4 technical signal
Google embodied reasoningGemini Robotics-ER 1.6 is described as a reasoning-first model for robotics, specializing in visual/spatial understanding, task planning, success detection, multi-view reasoning, instrument reading, and tool use.Blog dated 2026-04-14; available via Gemini API and Google AI Studio.Separates high-level embodied reasoning from low-level action/VLA control.🟢 Google DeepMindS4 developer-platform signal
Google failure-mode workflowGoogle invites developers to submit 10-50 labeled images illustrating specific failure modes if current capabilities are limited for specialized applications.10-50 labeled images; 2026-04-14 blog.Signals a developer feedback loop for robotics reasoning adaptation, but not commercial scale.🟢 Google DeepMindS3/S4 iteration signal
NVIDIA GR00T platformNVIDIA describes Isaac GR00T as an open reference platform for general-purpose humanoid robots, including open data / data pipelines, open robot foundation models, simulation frameworks, middleware, CUDA-X runtime libraries, and Jetson Thor.Developer page accessed 2026-06-13.Frames robotics AI as a full-stack platform rather than one model checkpoint.🟢 NVIDIA DeveloperS4 platform signal
NVIDIA training-data mixNVIDIA says GR00T models are trained on real captured data, synthetic data, and internet-scale video data, and can be post-trained for specific embodiments, tasks, and environments.Developer page accessed 2026-06-13.Makes the data recipe explicit: real + synthetic + internet video + post-training.🟢 NVIDIA DeveloperS4 data-stack signal
NVIDIA embodiment expansionNVIDIA says GR00T N1.6 training set expands embodiment coverage to Unitree G1, AgiBot Genie-1, Fourier GR-1, and specialized bimanual manipulation datasets.Developer page accessed 2026-06-13.Adds specific China / humanoid embodiment names to the model-stack evidence map.🟢 NVIDIA DeveloperS4 embodiment signal
GR00T-Dreams synthetic dataNVIDIA says Isaac GR00T-Dreams generates synthetic trajectory data from a single image and language prompt using Cosmos world foundation models.Technical blog dated 2025-06-16.Addresses the manual-demonstration bottleneck with a workflow claim rather than only model size.🟢 NVIDIA Technical BlogS4 synthetic-data signal
Synthetic-data time compressionNVIDIA says Research used GR00T-Dreams to generate synthetic training data for GR00T N1.5 in 36 hours, versus nearly three months using manual human data collection.36 hours vs nearly 3 months.Gives a quantified productivity anchor for synthetic data, not yet customer economics.🟢 NVIDIA Technical BlogS4 data-productivity signal
Open model checkpointHugging Face model card describes nvidia/GR00T-N1.5-3B as a 3B-parameter open foundation model for generalized humanoid robot reasoning and skills, ready for non-commercial use.3B parameters; BF16; safetensors; PyTorch; non-commercial use; version 1.5.Makes the model layer externally inspectable and developer-testable, with license limits.🟢 Hugging Face / NVIDIAS4 ecosystem signal
GR00T N1.5 architectureThe model card says GR00T N1.5 uses image frames, text instruction, and proprioception; it uses SigLip2, T5, and a flow-matching transformer / diffusion transformer for action chunks.Model card accessed 2026-06-13.Gives enough public architecture detail for technical due diligence; not economics.🟢 Hugging Face / NVIDIAS3/S4 technical signal

4. Interpretation: the new data-stack ladder

Add this ladder to the robotics field guide:

  1. Demo model: video shows a robot doing a task. S1/S2.
  2. Developer-accessible model or SDK: external developers can test / fine-tune / post-train. S3/S4.
  3. Data pipeline: teleoperation, simulation, synthetic trajectory generation, and post-training are documented. S4.
  4. Cross-embodiment adaptation: model/data stack works across multiple named robot bodies. S4 if source-backed.
  5. On-device / deployment runtime: model can run locally with latency / connectivity rationale. S4.
  6. Customer workflow proof: model improves uptime, intervention rate, task success, or deployment cost at customer site. S5 candidate.
  7. Financial proof: recurring software/model revenue, gross margin, attach rate, retention, or audited segment economics. S5+.

Current Google/NVIDIA evidence reaches steps 2-5 publicly. It does not yet reach steps 6-7 publicly.

5. Signal vs noise

Signal

  • A model is developer-accessible, downloadable, or fine-tunable with documented inputs / architecture / license. 🟢
  • A data workflow quantifies demonstration requirements or synthetic-data productivity, such as 50-100 demonstrations or 36 hours vs nearly three months. 🟢
  • On-device inference is tied to latency and connectivity constraints rather than generic “AI at the edge” marketing. 🟢
  • Cross-embodiment evidence names specific robot bodies such as ALOHA, Franka FR3, Apptronik Apollo, Unitree G1, AgiBot Genie-1, or Fourier GR-1. 🟢
  • Customer deployments later disclose model-attributable uptime, intervention-rate reduction, cycle-time improvement, safety-case compatibility, or ROI. 🟢 if future source.

Noise unless upgraded

  • “Foundation model for robots” without developer access, data recipe, embodiment coverage, or deployment metric. 🔴
  • “Synthetic data solves robotics” without real-world validation, task distribution, failure rates, or sim-to-real evidence. 🔴
  • “Open model” without license/commercial-use constraints. 🔴
  • “On-device model” without latency, compute, thermal, power, or reliability evidence. 🟠
  • “Partner ecosystem” without deployed robot counts, attach rates, contract value, or customer ROI. 🟠

6. Public-safe site draft section

The robotics AI stack is becoming a data factory, not a demo reel

The most important robotics AI update is not another viral humanoid clip. It is the emergence of a repeatable model-and-data stack.

Google DeepMind is pushing two separable layers. Gemini Robotics On-Device is a VLA model optimized to run locally on robots, with a fine-tuning path that Google says can adapt to new tasks with as few as 50-100 demonstrations. Gemini Robotics-ER 1.6 is a higher-level embodied reasoning model available through Gemini API and Google AI Studio, focused on spatial understanding, task planning, success detection, multi-view reasoning, instrument reading, and tool use.

NVIDIA is building a different but complementary stack. Isaac GR00T is positioned as an open reference platform for humanoid robots: open data and pipelines, robot foundation models, simulation frameworks, middleware, runtime libraries, and Jetson Thor. GR00T-Dreams uses Cosmos world foundation models to generate synthetic trajectory data; NVIDIA says this generated training data for GR00T N1.5 in 36 hours versus nearly three months of manual human data collection. The GR00T-N1.5-3B model card adds technical detail: a 3B-parameter humanoid model using vision, language, proprioception, and a flow-matching action transformer, available for non-commercial use.

The public implication is simple: robotics research should now track data-stack productivity alongside robot-body evidence. The best question is not “which robot demo looks best?” It is “which stack can turn fewer demonstrations, more simulation, and more real-world failure data into safer, cheaper, repeatable deployment?”

Footer: Evidence map only. No company ranking. No trade recommendation. Model/data-stack progress is S4 platform evidence until customer economics and software value capture are disclosed.

7. Slide-ready compression

Title: Robot AI is becoming a data stack, not a demo reel

Subtitle: Strong S4 platform evidence; not yet S5 commercial-economics proof.

Three-card layout:

  1. Google: local VLA + embodied reasoning

    • Gemini Robotics On-Device: local inference, fine-tuning, 50-100 demos.
    • Gemini Robotics-ER 1.6: spatial reasoning, planning, success detection, API access.
    • Source: Google DeepMind 🟢.
  2. NVIDIA: open model + synthetic data pipeline

    • Isaac GR00T: model, data, simulation, middleware, runtime, Jetson Thor.
    • GR00T-Dreams: 36 hours synthetic-data generation vs nearly 3 months manual collection.
    • GR00T-N1.5-3B: 3B parameters, non-commercial use, vision/language/proprioception/action model.
    • Source: NVIDIA / Hugging Face 🟢.
  3. What still blocks S5

    • No disclosed model ARR, attach rate, gross margin, customer ROI, uptime/intervention delta, or paid fleet economics.
    • Synthetic data and on-device inference are productivity signals, not proof of return cycle.

Footer: Evidence map only. No winner ranking. No trade recommendation. Data-stack progress ≠ scaled commercial economics.

8. What would change our mind

Upgrade signals

  • A robot OEM or customer discloses that Google / NVIDIA / another model stack reduced intervention rate, deployment time, or cycle time in a paid deployment. 🟢
  • Model provider discloses paid robot-model revenue, per-robot pricing, ARR, attach rate, gross margin, or retention. 🟢
  • Independent developers reproduce task adaptation with 50-100 demonstrations or lower across multiple robot embodiments and publish failure rates. 🟡/🟢 depending source.
  • Synthetic trajectory workflows show real-world task success lift across new environments, not only benchmark or demo gains. 🟢/🟡.
  • On-device VLA deployments disclose latency, power/thermal envelope, uptime, fallback behavior, and safety-case integration. 🟢.

Downgrade signals

  • Model stacks require large site-specific data collection for every new customer task, weakening scalability. 🟢/🟡 if disclosed.
  • Synthetic data improves demos but fails in contact-rich production environments. 🟢/🟡.
  • OEMs internalize the model stack, leaving horizontal providers with low value capture. 🟠 until revenue evidence appears.
  • Licensing limits or safety constraints slow commercial adoption despite open research access. 🟢/🟠.
  • Customer deployments prefer classical controls / task-specific automation because foundation-model reliability or certification is insufficient. 🟢/🟡.

9. Common misconceptions

  1. Misconception: “Robot foundation models mean robotics has already had its ChatGPT moment.”

    • Correction: public evidence shows platform formation and better data workflows, not yet scaled paid usage, gross margin, or customer ROI.
  2. Misconception: “Synthetic data removes the need for real robot data.”

    • Correction: NVIDIA’s own framing still combines real captured data, synthetic data, internet-scale video, teleoperation, simulation, and post-training. Synthetic data is a leverage mechanism, not a full substitute.
  3. Misconception: “On-device inference automatically means deployable robotics.”

    • Correction: on-device inference helps latency and connectivity, but deployment still needs reliability, safety cases, compute/power/thermal fit, customer workflow integration, and economics.
  4. Misconception: “An open model equals commercial value capture.”

    • Correction: an open / non-commercial model can accelerate ecosystem learning while leaving the business model unresolved.
  5. Misconception: “Cross-embodiment adaptation means one model can run any robot.”

    • Correction: named embodiment adaptation is meaningful, but every new robot/task/site still needs validation, safety analysis, and real-world failure tracking.

10. Think Deeper questions

  • Is the scarce asset in robotics the model architecture, the real-world failure dataset, the simulation/synthetic-data pipeline, or the customer deployment workflow?
  • If Google and NVIDIA commoditize parts of the model stack, where does value migrate: robot OEM, deployment integrator, compute supplier, data owner, or application software layer?
  • What is the best leading KPI for robot-model progress: demonstrations required, intervention-rate reduction, task success under distribution shift, model attach rate, or post-training cost?
  • Will on-device VLA inference become a safety / compliance requirement for latency-sensitive deployments, or only a developer convenience?
  • How should public research distinguish model benchmark progress from customer ROI progress?

11. Source list

12. Public-safety notes

PUBLIC-safe if used as a platform-evidence map. Do not include Hugo private portfolio data, watchlist logic, position weights, purchase prices, trade rationale, private channel checks, paid-report excerpts, or unverified rumors. Do not frame GOOG/GOOGL, NVDA, Apptronik, Unitree, AgiBot, Fourier, Google DeepMind, NVIDIA, or any robotics/security exposure as buy/sell/hold. Do not imply Google or NVIDIA has proven robot-model software economics unless future filings or primary customer disclosures provide revenue, margin, attach-rate, or ROI evidence.