Robotics tactile manipulation gate — when hands become labor, not demo
Date: 2026-06-17
As-of: 2026-06-17
Owner: Hugo / Genius Team
Research owner: Finance / Charlie AGT-002
Status: RESEARCH_ONLY
Visibility: PUBLIC
Output target: none; possible future /robotics/signals/, /robotics/dexterity/, or slide appendix only after Hugo review
Public-safety flag: PUBLIC-safe as an industry/framework evidence artifact. No Hugo private portfolio data, no trade recommendation, no private channel checks, no paid-report excerpts, no rumors.
0. One-line answer
Robotics 的最新高价值证据门不是“有没有五指手”,而是 tactile / in-hand / whole-body manipulation 是否能把一次性 demo 升级为可衡量的劳动替代能力;截至 2026-06-17,Figure Helix 02、NVIDIA GR00T N1.5 / reference humanoid、Sanctuary AI hydraulic-hand demos 已经给出更强的 S4 manipulation evidence,但仍不是 S5 economics,因为公开源仍缺 task trial counts、failure distributions、robot-hours、intervention rate、customer ROI/payback、repeat deployment、revenue / gross margin / service burden。🟢 primary sources; 🟠 Charlie S4/S5 synthesis.
1. Core question
如果 humanoid 的价值主张是替代真实劳动,那么“手”的证据应该如何升级?
Working answer: 需要把五个层级分开。
- Hand hardware: DOF、触觉、掌心摄像头、液压/电驱结构。🟢
- Skill demo: 拧瓶盖、取药片、打针筒、翻转方块、抓取金属件。🟢
- Measured manipulation: trial count、success rate、失败类型、reset / teleop boundary、object distribution。🟢/🟠
- Workcell / workflow proof: cycle time、shift duration、task count、robot-hours、uptime、intervention、safety。🟢/🟠
- Economics: ROI/payback、repeat order、service burden、vendor revenue/gross margin。🟢/🟠
当前公开证据明显从第 1–2 层推进到第 3 层边缘,Figure 的 whole-body demo 接近第 4 层展示;但没有完整第 4–5 层 denominator,所以仍应标 S4,不标 S5。🟠
2. Why this artifact is additive
已有 robotics-dexterity-benchmark-gate-v1.md 重点是 benchmark / measurement infrastructure:DexBench、RobOmni、Isaac Lab-Arena、GR00T FAQ、IEEE / safety standards。
本文件的增量是“把 benchmark gate 落到最新 public company / platform evidence 上”:
- Figure: tactile + palm cameras + whole-body controller,把手部动作与步态/平衡统一进 4 分钟连续任务。🟢
- NVIDIA: GR00T N1.5 / reference humanoid,把模型、触觉五指手、Jetson Thor、simulation / teleop / deployment workflow 接成开放 reference stack。🟢
- Sanctuary AI: hydraulic hand / tactile sensors / in-hand reorientation,强调接触丰富任务的 sim-to-real 和手部硬件路线。🟢
这不是“谁赢”的文件;它是 labor-substitution evidence gate。🟠
3. Evidence table
| Evidence anchor | Dated anchor | Quantified facts | What it proves | What it does not prove | Grade | Signal |
|---|---|---|---|---|---|---|
| Figure Helix 02 full-body autonomy | Figure AI official post, 2026-01-27 | 4-minute dishwasher task; 61 loco-manipulation actions; no teleoperation, no human intervention, no resets claimed; onboard sensors only. | Figure has shown a longer-horizon autonomous manipulation sequence that couples walking, balance and hands. | No disclosed repeated trial count, success-rate distribution, object variation, robot-hours, intervention over weeks, customer ROI/payback or revenue. | 🟢 | S4 autonomy/manipulation demo |
| Figure tactile + palm vision | Figure AI official post, 2026-01-27 | Fingertip tactile sensors detect forces as small as 3 grams; palm cameras support in-hand visual feedback; demos include unscrewing bottle cap, extracting pills, dispensing exactly 5 ml, singulating metal pieces from clutter. | Hand evidence is moving from vision-only grasping toward touch / in-hand feedback. | Does not disclose trial counts, failure cases, reset policy, durability, cycle time at customer site, or cost to deploy. | 🟢 | S4 tactile-manipulation signal |
| Figure System 0 whole-body control | Figure AI official post, 2026-01-27 | 10M-parameter neural prior; 1 kHz; trained on >1,000 hours retargeted human motion; 200,000+ parallel simulation environments; replaces 109,504 lines of hand-engineered C++. | Loco-manipulation is being attacked as a full-body control problem, not separate walk-then-grasp states. | Does not prove generalization across unseen homes/factories, safety certification, uptime, service burden, or economics. | 🟢 | S4 control-stack signal |
| NVIDIA GR00T N1.5 improvement | NVIDIA Research, 2025-06-11 | Language Table success: 52.8% N1 vs 93.2% N1.5; Sim GR-1 Language: 36.4% vs 54.4%; RoboCasa with 30 demos/task: 17.4 vs 47.5; DreamGen 12 new verbs: 38.3% N1.5 vs 13.1% N1. | Model progress is measurable on manipulation/language benchmarks and low-data tasks. | Benchmark improvements do not equal customer deployment; DreamGen “zero-shot” still trains on generated trajectories; real-world denominators remain missing. | 🟢 | S4 model-improvement signal |
| NVIDIA GR00T reference humanoid | NVIDIA Newsroom, 2026-05-31 | Unitree H2 body nearly 6 ft / 150 lb; 31 body DOF; dual Sharpa Wave tactile five-finger hands with 22 hand DOF; 75 total DOF; late-2026 availability from Unitree. | Research stack is becoming a standardized physical reference platform for dexterous humanoid work. | Late-2026 availability is not deployment proof; no customer economics, reliability, utilization, or safety outcomes. | 🟢 | S4 research-platform signal |
| NVIDIA GR00T workflow | NVIDIA Newsroom / Developer materials, 2026 | Isaac Teleop for demos; GR00T open models; Isaac Sim / Isaac Lab for simulate-train-test-evaluate; Isaac ROS / Jetson Thor for deployment / real-time inference. | Physical-AI tooling is becoming a full pipeline from data capture to on-robot inference. | Toolchain availability does not prove any one OEM’s robot can work economically in production. | 🟢 | S4 infrastructure signal |
| Sanctuary hydraulic hand disturbance test | Sanctuary AI official blog, 2026-04-01 | In-hand reorientation policy trained in simulation executed in real world; against gravity with 500 g added load not encountered during training. | Shows a high-contact sim-to-real manipulation result under an explicit disturbance. | Single demo does not provide trial count, full success distribution, task diversity, durability, labor ROI, or product economics. | 🟢 | S3/S4 dexterity hardware/control signal |
| Sanctuary letter-cube reorientation | Sanctuary AI official blog, 2026-04-01 | Target orientation achieved 10 consecutive times without dropping the cube; fingertip-only manipulation without palm support. | More concrete than a generic “dexterous hand” claim because it includes a repeated-count success anchor. | Still not a broad benchmark, customer workflow, uptime, safety or economics proof. | 🟢 | S4 in-hand manipulation signal |
| Sanctuary tactile sensors | BusinessWire / Sanctuary announcement, 2025-02-26 | Tactile sensors intended to support blind picking, slippage detection and prevention of excessive force; teleoperation pilots can use dexterity more effectively. | Touch is positioned as a practical enabler for occluded / fine manipulation. | Teleoperation-assisted capability is not autonomous labor economics; no intervention or support-cost denominator. | 🟢 company announcement via BusinessWire | S3/S4 tactile signal |
4. Signal vs noise
Signal
- Whole-body task sequences with duration, action count, autonomy boundary, reset boundary and sensor boundary disclosed. 🟢
- Tactile / palm-camera tasks where vision is occluded or contact force matters: pills, syringe, bottle cap, cluttered metal pieces, in-hand cube reorientation. 🟢
- Repeated-count claims such as 10 consecutive target orientations without dropping. 🟢
- Benchmark deltas with explicit baselines, success rates and data limits: e.g. N1 vs N1.5 on Language Table / Sim GR-1 / RoboCasa. 🟢
- Reference platforms that standardize body, hands, compute, teleop, simulation, training, evaluation and deployment. 🟢
- Future disclosure of task-level trial counts, failure distributions, resets, teleoperation boundary, robot-hours, uptime, intervention and safety incidents. 🟢 if disclosed.
Noise unless upgraded
- “Five-finger hand” without force/tactile, trial, durability or task denominator. 🔴/🟠
- Edited manipulation videos without reset, teleop, object-set and failure disclosure. 🔴/🟠
- A single successful object demo treated as general labor substitution. 🟠
- Simulation benchmark gains presented as customer ROI. 🟠
- “Zero-shot” claims without explaining whether generated trajectories, pretraining embodiments or task-specific data were used. 🟠
- Hardware DOF / compute specs presented as productivity proof. 🟢/🟠
5. Stage classification
Current label: S4 manipulation evidence, not S5 commercial economics.
Why S4:
- Figure reports a 4-minute / 61-action autonomous whole-body task with no teleoperation, no intervention and no resets. 🟢
- Figure discloses tactile sensitivity down to 3 grams and palm cameras, tied to specific manipulation tasks. 🟢
- NVIDIA discloses measured model improvements and an open reference robot with 75 total DOF, tactile five-finger hands and late-2026 availability. 🟢
- Sanctuary discloses 500 g disturbance transfer and 10 consecutive fingertip-only cube reorientations. 🟢
Why not S5:
- No source reviewed here discloses multi-week customer robot-hours, uptime distribution, intervention minutes, reset rate, safety incidents, maintenance/service hours, task ROI/payback, repeat order, vendor revenue, gross margin or service burden. 🟠
- Manipulation demos may be technically important while still being far from economically repeatable work. 🟠
- Reference platforms and model benchmarks can accelerate category learning but are not themselves customer-side acceptance evidence. 🟠
6. What would change our mind
Upgrade toward S5 if public primary/customer/filing sources disclose at least four of these:
- A standardized manipulation task suite with object-set size, trial count, success rate, failure taxonomy, reset policy and autonomy boundary. 🟢
- Real customer workflow results: cycle time, pieces handled, quality error rate, rework, safety events and human supervision per shift. 🟢
- Robot-hours over weeks/months with uptime and intervention distributions, not just a best-run video. 🟢
- Repeat deployment or renewal after measured acceptance in the same manipulation-heavy workflow. 🟢
- Service/maintenance burden for hands: sensor wear, finger/actuator replacement, calibration, failure rate and field support hours. 🟢
- Customer ROI/payback tied to labor capacity, throughput, quality, safety or downtime reduction. 🟢
- Vendor financial materiality: robot revenue, gross margin, RaaS economics, warranty/service cost, backlog/RPO or cash conversion. 🟢
Downgrade if:
- Tactile manipulation remains mostly single-object demos without repeated trial transparency. 🟠
- Whole-body demos do not transfer from kitchens/labs to paid customer workflows. 🟢/🟠
- Hand durability, calibration, sensor drift or actuator service burden becomes the limiting cost. 🟢/🟠
- Customers solve manipulation with fixed automation or task-specific end effectors instead of humanoid hands. 🟢/🟠
- Model gains plateau without lowering intervention rate or deployment time. 🟢/🟠
7. Public-safe site draft section
The next robotics proof is in the hand — but not because it has five fingers
A humanoid that can walk near a workstation is still not a worker. The harder proof is whether it can use contact: feel small forces, see from the palm when the head camera is occluded, adjust grip before slipping, recover when an object shifts, and keep doing this for many hours without an operator quietly saving the task.
The newest public evidence is better than the old demo cycle. Figure’s Helix 02 post claims a 4-minute dishwasher task with 61 loco-manipulation actions, no teleoperation, no human intervention and no resets. It also discloses fingertip tactile sensors sensitive to forces as small as 3 grams and palm cameras used for tasks such as pill extraction, 5 ml syringe dispensing and cluttered metal-part picking. That is a real S4 signal: whole-body autonomy and tactile manipulation are becoming measurable.
NVIDIA is pushing the same gate from the platform side. GR00T N1.5 reports higher benchmark success rates than N1, including 93.2% vs 52.8% on Language Table and 47.5 vs 17.4 on RoboCasa with 30 demos per task. Its reference humanoid combines a Unitree H2 body, dual Sharpa tactile five-finger hands, Jetson Thor compute and GR00T / Isaac workflows, with late-2026 availability from Unitree. This is not customer economics, but it makes dexterous humanoid experiments more standardized and reproducible.
Sanctuary AI adds a useful hardware/control contrast. It reports simulation-trained hydraulic-hand policies transferred to real-world in-hand reorientation with a 500 g added disturbance, and a separate letter-cube task achieving target orientation 10 consecutive times without dropping the cube. Again, this is strong manipulation evidence, not broad labor substitution.
The disciplined conclusion: hand evidence is improving from anatomy to task proof. But the S5 denominator is still missing: repeated trials, failure distributions, robot-hours, uptime, intervention, maintenance, ROI/payback, repeat deployment and financial materiality.
Footer: Evidence map only. No winner ranking. No trade recommendation. Five fingers are not labor substitution; tactile demos are not customer economics.
8. Common misconceptions
Misconception 1: “Five fingers mean human-level dexterity.”
Correction: fingers are hardware anatomy. Labor substitution needs task success across object variation, force control, failure recovery, durability, safety and economics. 🟠
Misconception 2: “A 4-minute autonomous demo proves the robot can work.”
Correction: it proves a stronger autonomy/manipulation milestone if the disclosed boundary is accurate; it does not prove multi-shift uptime, intervention rate, support burden, ROI or repeat deployment. 🟢/🟠
Misconception 3: “Benchmark gains equal customer value.”
Correction: benchmark gains are S4 model evidence. Customer value requires workflow outcomes and economics. 🟠
Misconception 4: “Touch sensors automatically solve manipulation.”
Correction: tactile sensing is an enabler; the evidence still needs task distribution, trial count, failure cases, calibration burden, durability and service cost. 🟢/🟠
Misconception 5: “Hydraulic / electric hand architecture decides the winner alone.”
Correction: hand architecture matters, but value capture depends on reliability, manufacturability, service burden, task fit, software/data loop and customer economics. 🟠
9. Think Deeper questions
- Which manipulation KPI should become the default public denominator: task success rate, intervention rate, reset rate, recovery success, or cost per completed task?
- Does value migrate to dexterous hands, tactile sensors, whole-body control models, teleop/data pipelines, simulation platforms, or deployment integrators?
- If tactile hands are high-service components, can humanoids reach acceptable gross margin before hand durability improves?
- Which customer workflow will first disclose manipulation ROI: logistics tote handling, kitting, machine tending, assembly, home chores, healthcare support, or retail backroom tasks?
- Should Hugo’s public robotics framework split “locomotion proof” and “manipulation proof” into separate S4 gates?
10. Source list
- Figure AI, “Introducing Helix 02: Full-Body Autonomy,” 2026-01-27. https://www.figure.ai/news/helix-02 🟢 primary company source; not independent validation.
- NVIDIA Research, “GR00T N1.5: An Improved Open Foundation Model for Generalist Humanoid Robots,” 2025-06-11. https://research.nvidia.com/labs/gear/gr00t-n1_5/ 🟢 primary research source; benchmark claims are not commercial proof.
- NVIDIA Newsroom, “NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research,” 2026-05-31. https://nvidianews.nvidia.com/news/nvidia-open-humanoid-robot-reference-design 🟢 primary company source.
- NVIDIA Developer, “Isaac GR00T - Generalist Robot 00 Technology,” accessed 2026-06-17. https://developer.nvidia.com/isaac/gr00t 🟢 primary product/developer source.
- Sanctuary AI, “Sanctuary AI Leads the Industry in Controlling Advanced Hydraulic Hands Using Reinforcement Learning,” 2026-04-01. https://www.sanctuary.ai/blog/sanctuary-ai-controlling-advanced-hydraulic-hands 🟢 primary company source; not independent validation.
- Sanctuary AI, “Sanctuary AI Demonstrates Zero-Shot In-Hand Manipulation on Hydraulic Hand,” 2026-04-01. https://www.sanctuary.ai/blog/in-hand-reorientation-policy-with-letter-cube 🟢 primary company source; not independent validation.
- BusinessWire / Sanctuary AI, “Sanctuary AI Equips General Purpose Robots with New Touch Sensors for Performing Highly Dexterous Tasks,” 2025-02-26. https://www.businesswire.com/news/home/20250226807548/en/Sanctuary-AI-Equips-General-Purpose-Robots-with-New-Touch-Sensors-for-Performing-Highly-Dexterous-Tasks 🟢 company announcement distribution; not independent validation.
- Charlie synthesis across tactile manipulation / S4-S5 evidence gates, 2026-06-17. 🟠
11. Public-safe boundary
Public-safe:
- Industry/framework evidence, source grades, dated anchors, S4/S5 stage classification, misconceptions, Think Deeper questions and source list above.
Exclude:
- Hugo private portfolio weights, purchase prices, tax context, private trade rationale, private channel checks, paid-report excerpts, unverified supplier/customer rumors, buy/sell/hold language, and any claim that Figure, NVIDIA, Sanctuary, Unitree, Sharpa, Tesla, Agility or any related public/private security is an investment recommendation.
Do not imply:
- Figure’s demo proves home or factory economics.
- NVIDIA’s reference robot proves late-2026 deployment economics.
- Sanctuary’s hand demos prove full humanoid labor replacement.
- Any benchmark, reference hardware, tactile hand, or model improvement equals S5 commercial proof without the customer/economic denominator.