【請求書・インボイス→銀行振込支払対応のお知らせ】 5,000円以上のレポートは、当サイト発行の請求書(適格請求書/インボイス)による銀行振込に対応いたしました。
お問合せ:info@sattu-ai-agent.com

Valuable Japanese physical AI practices: Physical AI Day (15 October 2026): a briefing from public materials

Read from the public record: where Japanese teams say the stack actually breaks

Introduction

This briefing is a pre-event reading of Physical AI Day (15 October 2026, Tokyo), hosted by CoreWeave and Weights & Biases with NVIDIA as a partner. It is assembled only from the official session abstracts and from materials the speakers and their organizations had already put in public: press pages, IR, program results, product sites, and earlier talks. There are no slides and no recordings yet. Treat it as a map of claims, not as minutes.

The day’s stated purpose is unusual for a vendor-adjacent event. Physical AI fails in pipelines—collection, annotation, evaluation, Sim2Real, deployment—more often than it fails in a named model. Those failures rarely become case studies, so separate teams rerun the same expensive loops. The agenda is built to surface difficulty: multi-finger mechanisms, trajectory-level RL, dual-arm handover, used-goods inspection, opening industrial controllers, and why training and simulation do not fit a fixed on-prem box.

The seven main talks form one stack when read together. NVIDIA sketches the whole-system design (VLA and WAM plus conventional programmable control). Weights & Biases and CoreWeave treat the learning unit as a trajectory and insist that metrics, video, and traces live on the same run. Algomatic Dynamics argues that multi-finger success is decided by the hand and the collection loop, not by model scale. ABEJA reports putting foundation models onto the open humanoid-style arm OpenArm. Mercari takes the same class of limits into a warehouse task: shoe inspection for cross-border trade. FANUC’s move is not to win a foundation-model race but to open more than a million installed industrial robots over ROS 2, Python, and 1 ms external command. AWS Japan’s claim is narrower than it sounds: millisecond control stays on the robot; burst training, synthetic data, and parallel simulation do not.

For engineers already hitting hardware and data walls—dexterous hands, contact-rich dual-arm work, viewpoint shift, sparse rewards, factory interfaces—the value is not a list of Japanese product names. It is a public vocabulary for the same seams: which layer each organization has decided to own, which numbers are still targets, and which problems they are willing to name before the event. Nothing here is Tesla-internal practice, and nothing here is a procedure for mass-producing a humanoid. It is what one Japanese program was willing to write down in advance about the loop that sits under the demo.

Event page: https://wandb.ai/site/resources/events/physical-ai-day/

Hosts: CoreWeave and Weights & Biases, with NVIDIA as a partner.
When / where: 15 October 2026, 13:00–18:30 JST, Kioi Conference, Tokyo Garden Terrace Kioicho.

The organizers’ stated purpose is that Physical AI is hard in model training, simulation, data collection, evaluation, and deployment, while “what failed,” “where it hurt,” and “what you only learn by doing it” rarely make it into blogs. Teams therefore repeat the same experiments in isolation. The day is meant to share difficulty as well as success. Later sessions include NVIDIA, AWS, and Weights & Biases tools, but the frame is not a vendor lecture: vendors are also supposed to learn the developers’ constraints.


NVIDIA

Ken Arai
Solutions Architecture & Engineering
Senior Manager, Robotics Developer Relations
Approaches and outlook toward putting Physical AI into practice (13:05–13:25)

Official abstract

Putting Physical AI into practice requires a whole-system design: robot learning methods such as VLA and WAM, combined with conventional programmable modeling. The talk will organize supporting problems—data-pipeline design, management and quality of training data, benchmarking—and present a practical approach that connects training, evaluation, and deployment on real robots, plus an outlook.

What the public materials imply

This is the opening 20 minutes. It is a map of the stack, not a success story on one task. The abstract names two learning families first:

  • VLA (Vision-Language-Action): a learned model that turns vision and language into action. ABEJA, Mercari, and AWS use the same term later.
  • WAM: listed next to VLA. In this agenda it stands as the other robot-learning family, set against the older habit of writing the whole stack as programs.

The abstract’s real claim is that a learned model alone does not ship. The design problem is how to combine learning with programmable modeling: teach pendants, trajectory planning, force control, safety PLCs, classical vision.

Four development problems named in the abstract

  1. Data-pipeline design
    Who collects, with which cameras, at which rate, into which action space. Collection methods split later: leader–follower, VR, UMI, conversion of human work.
  2. Management and quality of training data
    Volume is not enough. Failed trajectories, contact, viewpoint shift, and missing force/tactile signals are quality problems. This is the same point as the W&B talk: keep video and traces on the same run.
  3. Benchmarking
    “Model A beats model B” is meaningless unless robot, lighting, and cameras are held fixed. That is one reason OpenArm-style evaluation cells get standardized.
  4. Connecting train → evaluate → deploy
    NVIDIA’s public stack here is Isaac (simulation / learning), Cosmos (world models), and GR00T (robot foundation models). Sim2Real—whether a policy that runs in the simulator still runs on an industrial arm or a humanoid-style arm—is the boundary of “in practice.”

How this talk feeds the rest of the day

  • VLA → ABEJA dual-arm work, Mercari shoe inspection, AWS training infrastructure
  • Coexistence with programmable control → FANUC teach language vs ROS 2 / Python
  • Data quality and evaluation → W&B experiment tracking, Algomatic collection pipeline
  • Simulation → FANUC ROBOGUIDE × Isaac Sim, AWS G6e / G7e

Twenty minutes is enough to list the seams in the pipeline, not to reproduce papers. Which product generation of Isaac / Cosmos / GR00T appears on the day is not in the public abstract.


CoreWeave · Weights & Biases

Yuya Yamamoto
Senior Account Solutions Architect
Trajectory-level reinforcement learning and experiment management in Physical AI (13:25–13:45)

Official abstract

Reinforcement learning of VLA models in Physical AI requires evaluation and management of complex trajectories that interact with the environment. The talk will organize how to apply GRPO—critic-free, trajectory-level updates—in an action-token variant and a flow-matching variant, and will present experiment management that unifies numeric metrics with video and traces. It will also share benefits seen in practice and problems such as advantage collapse.

Technical content visible from public sources

After imitation learning, if you add RL, this talk treats the unit of optimization as the trajectory, not the step (token or control tick).

GRPO (Group Relative Policy Optimization) does not train a separate value function (critic). It samples several trajectories under the same condition and takes advantage from their relative quality. The public reading is that a pattern popularized in LLM reasoning training (DeepSeek-R1 and related work) is being moved onto VLA action sequences.

Two variants named in the abstract:

  • Action-token type: actions as a discrete / token sequence; updates from relative reward inside the group.
  • Flow-matching type: for VLAs that generate continuous actions with flow matching (π0-class and similar). Same GRPO idea, adapted to a different generator.

Advantage collapse is named in the abstract. If reward differences inside the group flatten, relative advantage goes to near zero and updates stop (or become noise). Physical AI rewards are sparse in characteristic ways—every trajectory succeeds, every trajectory fails, only length differs, contact credit appears only at the end—and that sparsity makes collapse more likely. That is the public problem statement.

Later field talks explain why trajectory-level updates exist. Dual-arm handover does not make “which instant was wrong” obvious. Force in shoe inspection does not show up in step loss. Cable tension is a whole-trajectory score. Training a stable step-level PPO critic on robots is expensive; critic-free GRPO is the practical candidate in this framing.

Experiment management

W&B’s claim is that a Physical AI experiment cannot be reproduced from a loss curve. The abstract says a run should keep:

  • numeric metrics
  • video (rollouts)
  • traces (trajectories, sensors, commands)

CoreWeave is compute; W&B is the reproducibility layer. Trajectory-level RL multiplies rollouts, so compute and logs swell together. That is why the two firms share a slot.

The abstract promises both benefits and failure modes, which matches the event’s “share the pain” brief. Concrete VLA names, reward design, and group size are unpublished.


Algomatic Dynamics, Inc.

Yuki Nanri
Representative Director and CEO
The difficulty of developing multi-finger hands (13:45–14:10)

Official abstract

Task success on a multi-finger hand is not decided by the model alone. Mechanism and accuracy of the hand, teleoperated data collection, annotation, and training: if any stage in that end-to-end pipeline becomes the bottleneck, success rate falls. From the experience of modifying an open-source hand and running collection through training in-house, the talk will say what actually mattered at each stage and what became the bottleneck.

Speaker and company, from public sources

Nanri is CEO of Algomatic Dynamics. He is a Keio graduate and a research staff member in Kenji Tanaka’s lab at the University of Tokyo. After serving as CTO of Algomatic’s generative-AI business, he founded Algomatic Dynamics on 27 April 2026 as a robotics / Physical AI company. The mission is “Bringing smoothness to every robot.” In a company note he defines that as smoothness, not speed, strength, or cleverness.

On 9 September 2026 the company announced a ¥5 billion raise with DMM.com as subscriber, and a plan to launch an AI multi-finger hand platform in Japan within 2026. Public business lines:

  • turn video of skilled work into motion data and train with cross-embodiment; robot rental and operators on dispatch
  • sell AI multi-finger hands as hardware plus software, with task fitting, retuning, and maintenance
  • give characters a body using techniques from animatronics

The published technical claim is a contact-rich multi-finger hand plus one platform from collection → additional training of a foundation model → deployment on hardware. Targets run from rigid industrial parts to irregular food.

On X shortly after founding he wrote that they were building an imitation-learning package for a multi-finger hand, iterating on collection methods, that CG collection felt high quality, that software was approaching the ideal, and that the real robot still needed more tuning. That public post and the official “we ran one full loop” abstract are very likely the same work.

Bottlenecks as the abstract cuts them

Success rate is declared not to be a model-only number. Four seams:

  1. Hand mechanism and accuracy
    Joint count, wiring, backlash, stiffness, durability, and sensor placement act at once. If repeatability is poor, the same demo scatters.
  2. Teleoperated collection
    Leader–follower, gloves, collection via CG. Nanri has said publicly that collection methods are still trial and error.
  3. Annotation
    Contact, grasp success, force, object state. Video alone does not become training signal.
  4. Learning
    If the first three stages dirty the data, a larger model does not raise success.

The primary source for the talk is “we modified an OSS hand and ran collection through training ourselves.” This session covers work that off-the-shelf grippers and OpenArm’s default parallel gripper do not do well: soft objects, tools, buttons, bags, food. Finger count, sensor layout, and success rates are not in the public abstract.

This is the last talk before the break. Arms, inspection, and industrial robots in the second half sit on top of this “hand plus data” loop.


ABEJA, Inc.

Fumiaki Iwaki
Embodied Intelligence Group / Data Scientist
Trial and error in applying foundation models to a humanoid-style robot arm (14:20–14:45)

Official abstract

Social deployment of Physical AI turns on how evolving robot foundation models are applied to robots in the real world. The talk will present ABEJA’s work: trial and error in realizing dual-arm coordination with a VLA on the humanoid-style arm OpenArm, and the use of reinforcement learning. It will include an outlook: where practice is now, and what it may become.

Iwaki’s personal biography is almost absent from the public web. Title as above. If Nanri’s talk is “we ran a hand and a collection pipeline once around,” this talk is trial and error on the other side: putting existing foundation models onto an open humanoid-style arm.

OpenArm, the platform named in the abstract

OpenArm is a fully open-source 7-DoF humanoid arm from Enactic in Tokyo, designed for Physical AI research.

Public specs, in brief:

  • 7 axes per arm, human scale (about 160–165 cm stature)
  • rated payload 4.1 kg / peak 6.0 kg
  • QDD (quasi-direct drive), backdrivable
  • leader–follower, bilateral force feedback
  • CAD, firmware, ROS 2, Isaac Lab, MuJoCo all public
  • a dual-arm kit is relatively cheap for research (public price band around $6,500)

“Humanoid-style robot arm” here is not a full-body humanoid. It is an upper-body dual-arm OpenArm. It is becoming a default research platform on which contact-rich work, teleop collection, imitation, and real-robot evaluation can all be run.

Internal ABEJA slides already list SO-101, OpenArm Mini, OpenArm KER, OpenArm leader/follower, and VR controllers as collection tools, then send the demos into π0.5 or SmolVLA. The experiment rig for this talk is plausibly that same line.

What “trial and error” means in the abstract

Three points:

  1. Putting diverse foundation models onto a real robot
    Not a model-only benchmark. Application to one body, OpenArm. Camera placement, joint representation, action space, and control rate disagree here.
  2. Dual-arm coordination as a VLA
    Harder than single-arm pick-and-place. Handover, transfer, reorientation, two hands on one object.
  3. Use of RL
    Demos are not enough for dual-arm contact work; RL is added downstream. That connects technically to the W&B talk on trajectory-level GRPO.

The wording is “trial and error,” “current location,” “future picture,” not “we succeeded.” Expect seams, not a headline success rate.

ABEJA’s public record against this talk

ABEJA is listed (TSE 5574). The core business is ABEJA Platform for mission-critical AI operations. Physical AI is written into IR materials as the pillar after LLMs. Recent public projects line up with the abstract:

DateWhatRelation to this talk
Mar 2025Full member of AIRoArobot data ecosystem
Jul 2026NEDO project: factory visuo-tactile data → VTLA foundation model (Kawasaki, FANUC, Yaskawa, Osaka Univ., FingerVision)vision plus touch
Aug 2026Joint validation with Murata Manufacturing: VLA dual-arm part handover succeeded in a validation setupdual-arm coordination itself
Sep 2026JR East, TAKANAWA GATEWAY CITY: dual-arm recognize / grasp / load baggagefield-adjacent dual-arm work
Sep 2026Acquisition, Technology & Logistics Agency: VLA on UGVs and “VLAOps” requirementsoperations layer

The Murata case is almost the same task as the abstract. A VLA infers from camera images and completes “grasp with left → hand to right → reorient → insert into a rack” on hardware. The published caveat is “in a validation environment.” This session is the natural place to talk about the failures, the tuning, and OpenArm-specific problems behind that demo.

ABEJA’s repeated term is VLAOps: do not run the model once and stop; build a pattern for additional training and continuous operation. The IR story is to carry LLM operations practice onto robots.

Where dual-arm VLA tends to jam, from public materials

Not the talk itself; inferred from the abstract, ABEJA slides, and OpenArm’s design.

  • Collection cost. Dual-arm leader–follower is expensive. ABEJA materials compare VR teleop (Quest 3S) at roughly half the cost.
  • Gravity compensation and operability. One OpenArm is about 5.5 kg. Bad compensation dirties demos.
  • Camera vs action-space mismatch. Wrist cam, head cam, joint-command dimension and rate differ by model. π0.5 and SmolVLA split the preprocessing.
  • Credit assignment across arms. Handover does not say which arm was wrong. That is why trajectory-level RL appears.
  • Contact and compliance. OpenArm is backdrivable and therefore safer; a policy trained on a stiff industrial arm will put out force differently.

Twenty-five minutes favors a list of settings that worked and settings that did not, not paper reproduction. Iwaki’s papers are not findable in public search. Which VLA is used is not in the abstract. Whether the talk is the Murata task or another OpenArm task is unknown. Algorithm and sim-vs-real for the RL piece are unpublished.


Mercari, Inc.

Norimasa Kobori
Head of Research, R4D
R&D on automating inspection work with AI robots (14:45–15:10)

Official abstract

As cross-border trade grows, more goods bound for overseas shipping are collected first in Mercari’s domestic warehouses for inspection. Inspection is one of the processes where putting a robot in has the largest effect. This project tackled automating shoe inspection. Current models that drive robots have three limits: force is hard to control, accuracy drops when the camera moves, and training data is hard to collect. The talk presents Mercari’s work on each.

If ABEJA is “putting a VLA on an open dual-arm,” this is how those limits show up in a real warehouse task (shoe inspection). It is the first session whose subject is a C2C operations task.

Speaker

Norimasa Kobori is Head of Research at Mercari R4D, in post since September 2025.

Public background:

  • master’s work at Waseda’s humanoid robotics lab
  • Sony, Toyota Motor, Toyota Motor Europe (image sensors, robotics, automated driving, overseas research leadership)
  • doctorate from Nagoya University
  • Principal Researcher at Woven by Toyota / Woven Alpha, work around Woven City AI
  • left Toyota after about 20 years at the end of August 2025 and joined Mercari

This is an operations talk given by a research head who has shipped real-world robotics, not only a research-manager slot.

Why Mercari is robotizing inspection

The published business reason is simple.

  • September 2025: launch of the worldwide Mercari Global App
  • Cross-border volume has risen sharply in recent years (AWS results briefing: “19× in four years,” target of 50+ countries by 2028)
  • Overseas shipping requires inspection that photos and listing text match
  • SKUs are mostly unique used goods, not catalog parts; apparel and bags are soft; buttons and zippers appear
  • Classical model-based control does not automate this easily

R4D’s public definition of robotics matches: replace C2C physical steps—unpack, inspect, authenticate, photograph, pack, ship—with robots. Shoes were chosen as a high-step, high-leverage process. The AWS results briefing describes shoe inspection as nine steps.

At that briefing Kobori said the target is 20% automation of Mercari warehouses by 2028 and ¥200 million per year in cost reduction. That is a development report with a business KPI, not a lab talk.

The three limits in the abstract, and the public countermeasures

The three limits match R4D’s research themes and the GENIAC award text.

Limit in the abstractPublic work
Force is hard to controlExtend to a multimodal VLA with touch and force. Counter to vision-only policies that handle roughly
Accuracy drops when the camera movesCamera-Aware VLA. Train-time vs infer-time viewpoint. Synthetic images used to fill infer-time views, in public descriptions
Training data is hard to collectA collection stack that transfers human work feel onto the robot. AWS results briefing cites UMI (Universal Manipulation Interface) as making collection easier

GENIAC selection (9 September 2026) puts the pilot in Mercari inspection and fulfillment sites and funds the three threads as a national project. The session title says “inspection”; the award theme is “listing and inspection,” so photography and match-to-listing text may be in scope.

AWS Japan’s Physical AI support program (selected March 2026, results late August) already tested these three issues on shoe inspection. The October talk is that line of work, given just after the GENIAC award.

Technical stack as Mercari names it in public

  1. Task: shoe inspection, many steps, soft and irregular objects, condition checks.
  2. Base model: VLA from video and language to action.
  3. Extension: multimodal VLA with touch and force. Same direction as VTLA (Vision-Tactile-Language-Action) in the manufacturing project ABEJA joined.
  4. Viewpoint robustness: Camera-Aware VLA. Warehouses do not hold camera pose fixed.
  5. Data: not only teleop; convert human work into robot training. UMI is one example.

Used shoes are not a new-parts line: dirt, collapsed shape, laces, insoles, presence or absence of a box change every time. Generalization is the point. Robot make (dual-arm vendor, hand) is unpublished. Success rate and throughput are unpublished. UMI and Camera-Aware VLA are described only at results-briefing level. “20% of warehouses / ¥200 million a year by 2028” is a target; current attainment is unknown.


FANUC Corporation

Masahiro Morioka
Chief Engineer, Robot Mechanism R&D Division
Robot R&D Group
FANUC’s open-platform strategy and Physical AI in practice (15:10–15:35)

Official abstract

FANUC robots, with a cumulative installed base of more than one million units on production floors worldwide, accelerate social deployment of Physical AI through open-platform support such as ROS and Python. The talk will present practice of “see, think, move” Physical AI automation through the latest robot technology: a broad, reliable lineup, the open platform, and a high-accuracy digital twin.

This is an industrial OEM moving from a closed teach culture toward letting outside AI drive the robot over ROS 2 / Python / high-rate external command. If Mercari is “shoe inspection in our warehouse,” FANUC is “make the robots already in the factory into a Physical AI runtime.”

FANUC’s own definition is explicit. Physical AI is “AI that acts autonomously not only in virtual space but in physical space.” The company’s role is not primarily to invent the model, but to be a trustworthy, open robot substrate. Cumulative robot shipments passed one million in 2023.

Speaker

Masahiro Morioka is Chief Engineer of the Robot Mechanism R&D Division.

Public career:

  • March 1999: M.Eng., precision machinery, University of Tokyo
  • April 1999: joined FANUC, mechanism design in the robot lab
  • 2011: development manager, robot lab
  • 2019–: Chief Engineer, Robot Mechanism Development Lab (now Robot Mechanism R&D Division)

The point is that a mechanism lead is speaking—not software comms. He has designed machines from 3 kg payload to 2.3 t. He also spoke in December 2025 to students on “robots that support Japanese manufacturing.”

Three pillars of the open platform

Fixed in official messaging since around the December 2025 International Robot Exhibition:

PillarContentIntent
Official ROS 2 driverPublic on GitHub, ros2_control, 1 ms cycleconnect research and startup software straight to the machine
Native Python on the controllerPython runs on the controller; no PC hop requiredput the AI team’s language on the factory controller
Stream Motionposition, velocity, torque commanded externally at 1 msstream externally generated trajectories onto the machine smoothly

The claim is repeated: not only the CRX cobot, but the line from 3 kg class to 2.3 t. Not “we opened a lab arm.” The same interface on the factory workhorse.

CEO Kenji Yamaguchi, in a February 2026 interview, described the same three points as the way to accelerate Physical AI implementation. Company policy, not a side experiment.

Public partnerships

The special site first named two partners; a Fujitsu group was added later.

NVIDIA

  • FANUC robots as OpenUSD SimReady assets in Isaac Sim
  • real-time link between FANUC’s own ROBOGUIDE and Isaac Sim, reproducing the same trajectories and cycle times in a virtual plant
  • edge builds on Jetson Thor
  • official language: demonstration of imitation learning with Isaac GR00T has started

This is the industrial-robot side of Isaac / Cosmos / GR00T from Arai’s opening talk.

Google

  • May 2026: Physical AI robot system using Gemini Enterprise
  • agent understands instructions, recognizes objects, drives multiple robots
  • September 2026: “AI welding agent.” Reads drawings, generates current, voltage, and robot motion. Marketed as zero setup / zero teaching. Shipments from end of December
  • sold as a subscription on existing FANUC contracts—a field-installation shape, not a research demo

This is productization on welding, an existing application, not a lab sketch.

Fujitsu + Yaskawa + Kawasaki Heavy Industries

  • July 2026: agreement to study a Physical AI business
  • concept: Fujitsu Kozuchi Physical OS as an open substrate
  • NVIDIA technology in the mix
  • reporting: implement at Fujitsu’s Kahoku plant by end of September 2026, supply to the three firms by December

FANUC is also in the NEDO visuo-tactile data project that ABEJA joined. A major industrial OEM is showing up on both the foundation-model side and the control-OS side.

Official demo list

The open-platform case page is the likely slide source.

  1. Voice command; a generative model writes and runs Python (multiple languages)
  2. Detect a person, retreat without stopping the job, resume the path when they leave
  3. Dual-arm routing of a soft cable while watching tension
  4. Chase a moving part and drive screws
  5. AI-agent kitting, shown from the robot show onward
  6. Grape thinning with University of Yamanashi (agriculture)

The common move is to replace offline teaching with language, vision, and contact. Dice-face sorting and stacking on a colored block appear in the CEO interview as generative-AI examples.

Twenty-five minutes likely runs: three pillars, NVIDIA / Google demos, productization such as the welding agent. A mechanism chief engineer may also talk payload, accuracy, and safety. “Zero teaching” is marketing language for the welding agent; field yield is unpublished. GR00T imitation is at “demonstration started.” Physical OS with Fujitsu is a business study plus a schedule; APIs and responsibility splits are not visible. How TP programs coexist on the floor with Python / ROS 2 is not in the official text.


Amazon Web Services Japan G.K.

Yoshitaka Haribara
Principal Startup Solutions Architect
Why Physical AI development needs the cloud (16:15–16:35)

On the official agenda this slot follows a 15:35–15:45 break and a 15:45–16:15 panel.

Official abstract

Physical AI development needs a wide stack: collection and preprocessing pipelines, training and fine-tuning of foundation models such as VLA and WAM, simulation and Sim2Real, and deploy / orchestration to the edge. This session shows how AWS cloud services can make that work efficient, with examples from the Physical AI development support program.

The title is the claim. Not models or bodies: learning, simulation, and data expansion on prem alone is not realistic. If FANUC “opens the robot,” AWS is “where the compute and data sit before that.” Pairing VLA and WAM in the abstract uses the same vocabulary as Arai’s opening talk.

Speaker

Yoshitaka Haribara. Public titles vary by date; the event page says Principal Startup Solutions Architect; other materials add Frontier AI / Senior. Visiting associate professor at Osaka University’s Center for Quantum Information and Quantum Biology.

Public background:

  • B.S. mathematics, Osaka University; Ph.D., University of Tokyo Graduate School of Information Science and Technology (quantum optics / combinatorial optimization)
  • joined AWS Japan as a new graduate in 2018
  • generative-AI startup support; launched the AWS LLM Development Support Program in 2023
  • support for GENIAC round-2 awardees
  • launched “Physical AI Development Support Program by AWS Japan” in January 2026
  • author of a practical guide to building generative-AI apps on AWS
  • X: @_hariby

The person who ran Japan’s early LLM support program is applying the same pattern to Physical AI. He also gave the results-briefing wrap-up.

Why AWS says the cloud is required

The official reference cuts development as a three-step loop:

  1. Data generation and collection
    Real-robot demos are not enough. Parallel Isaac Sim (and similar), synthetic data, expansion of teleop data.
  2. Model training
    Fine-tunes and large training runs of VLAs / world models. One GPU is not enough; a cluster is.
  3. Serve and infer
    Push trained models to the edge (robots, Greengrass) and pull field data back to the cloud.

Public reasons for putting that loop on the cloud:

  • Bursty compute. Training and simulation are not always at full load. Renting hundreds of GPUs when needed is faster than a fixed on-prem fleet.
  • Filling the data hole. FastLabel expanded 90 teleop patterns about 100×, dropped bad samples, and reached 10,000 patterns. Average task success rose 14.2%. That scale is not collected on hardware alone.
  • Parallel sim. Study sessions recommend G6e / G7e (L40S or RTX PRO 6000 Blackwell) for rendering.
  • Training clusters. SageMaker HyperPod, AWS ParallelCluster, Capacity Blocks for ML. Highlanders was presented as training at the 100-billion-parameter class for 24 hours on ParallelCluster.
  • Edge loop. IoT Greengrass for model delivery; S3 / FSx for Lustre for data; then back into training. Control stays on site; the improvement loop sits in the cloud.

The core of the title is a rebuttal of “Physical AI = all edge.” Millisecond control on site; learning and data expansion in the cloud. That is the official split.

Haribara’s wrap-up at the results briefing: training foundation models, collecting and preprocessing large data, and running hundreds of simulations in parallel on prem alone is not realistic in time or cost. Being able to take the required compute at the required scale is not one option among many; it is indispensable.

What the program actually did

Started 27 January 2026. 51 companies selected (more than half startups). Credits up to $6 million in total. Support roughly March–August. Results briefing: 13 projects for press, 26 overall.

Overlap with this event:

  • Mercari was a selected company. Distributed training on AWS; nine-step shoe inspection; 20% warehouse automation and ¥200 million / year by 2028 as the target
  • NVIDIA study sessions at AWS Meguro on Isaac / GR00T and instance choice for data generation
  • Physical AI Scaffolding Kit (PASK) published: sample containers to fine-tune GR00T or π0 on SageMaker HyperPod

Haribara’s talk is therefore not abstract. It redraws Mercari and the Isaac stack FANUC is connecting, from the cloud side.

Examples where the results briefing said compute scale mattered:

CompanyWhat they ran in the cloud
Telexistencemulti-GPU distributed training. 10B params, 86% success (80% with public pretraining, 50% with none)
FastLabel90 → 10,000 patterns via synthetic expansion. +14.2% success
Highlanderslarge training kept running on ParallelCluster
MW50,000 hours / year collected; daily S3 → HyperPod loop; 2.5 PB target
Daifuku × JDSCunseen-item sort in simulation, 97 / 100
Ricohshared representation across 3 sim embodiments and 1 real; held LIBERO-10 performance at 1/3 model size
Mamezoualignment success 85% vertical / 80% horizontal; patterns cut from 2,000+ to 43

Other names at the briefing: ATOM (walk and sort; 24-hour continuous demo on 28 August; future target 200 robots, 300,000 hours, tens of billions of parameters), OMRON SINIC X (three patents / papers in three months), ZEALS (10× data conversion, hospital demo, 2026 target 100 robots / 10,000 hours), Takenaka, NRI.

Reference architecture

Haribara’s April 2026 GENIAC deep dive is the likely slide ancestor.

  • real robot / simulator → S3 / FSx
  • training: SageMaker HyperPod or ParallelCluster
  • simulation: EC2 G7e and similar
  • inference delivery: SageMaker Inference and IoT Greengrass
  • developer desk: Remote AWS Develop Station (remote desktop on GPU)

PASK is a starter kit so teams fine-tune a public VLA on HyperPod before they build a foundation model from scratch.

The through-line of his career is easy to tell in 20 minutes: LLM support program (2023) → GENIAC support → Physical AI program (2026). At the results briefing he said future support folds into the generative-AI practical-deployment program.

“Cloud is required” is about training and simulation. Official materials do not say to put the control loop itself in the cloud. The $6 million in credits is for the whole program; per-company amounts are unpublished. PASK and the reference architecture are samples; factory safety certification and air-gapped networks are separate problems. Success rates at the briefing are self-reported; there is no common benchmark.


What the seven public sessions say when read together

  1. The loop, not the model name.
    NVIDIA and AWS both put data → train → real robot → retrain ahead of a single score. Arai’s pipeline / quality / benchmark list and Haribara’s three-step cycle are two ends of one picture.
  2. Contact and two hands are the current working theme.
    Nanri’s multi-finger hand, Iwaki’s handover, Kobori’s shoes, FANUC’s cable routing all sit outside single-arm imitation with a parallel gripper.
  3. That is why trajectory-level learning and logging appear.
    Dual-arm credit, force, viewpoint change do not show in step loss. Yamamoto’s GRPO and advantage collapse, and W&B’s video-plus-trace, are the observation layer for that.
  4. The factory side is in an “open the interface” phase.
    FANUC’s ROS 2 / Python / 1 ms command is the interface that lets a research OpenArm policy or a VLA land on installed plant equipment. The installed base cited is one million-plus robots.
  5. Cloud is not a substitute for control.
    Learning, synthetic data, and clusters. AWS and FANUC do not contradict each other on that cut.
  6. The vocabulary is aligned across the day.
    VLA and WAM appear as a pair in both Arai’s and Haribara’s abstracts. OpenArm, GR00T, π0 / SmolVLA, Isaac Sim, UMI, Camera-Aware VLA, VLAOps, and Stream Motion name one stack from different chairs.
  7. Most published numbers are still targets or self-reports.
    Success in a validation cell, 20% of warehouses by 2028, welding-agent ship date, program success rates. Field yield is what to watch for in caveats on the day.

How a US robotics team can use this without over-reading it

Public Japanese practice and a US humanoid program (including a company scaling a multi-finger, contact-rich upper body) rhyme in a few places that are already on the record:

  • Hands and collection pipelines set the ceiling before the policy does. That is Algomatic’s official sentence, and it is the same class of constraint discussed around dexterous humanoid hands in the US.
  • Dual-arm handover and force-rich inspection need trajectory-level credit and logs, not only step loss.
  • Opening a 1 ms external command path on a reliable body is a product decision, not a model decision. FANUC is making that decision on an installed industrial base; a humanoid line has to make an analogous decision on its own controllers.
  • Burst training and synthetic expansion live in the cloud; the millisecond loop does not.

None of that is a transfer of unpublished Optimus hardware practice. It is a map of what Japanese teams are willing to say in public about the same layers.

What exists today is abstracts plus press pages, results briefings, IR, and earlier talks. Which demo, which model name, and how much failure each speaker will actually show will only be known from the day-of materials. The 15:45–16:15 panel has no official abstract. This note is preparation. It is not minutes.

さっつーのよい知らせ:最新話

【さっつーのよい知らせ】第19話・仲直り(アナログイラスト・漫画×癒し)

【漫画×アニメ】サメじろう家族の結末。仲直りか、それとも…|さっつーのよい知らせ・19話 仲直り

【あらすじ】

「ーー確かに、パパさんのやってきたことは良いこととは言えないけど、それでも、ゆるしてあげよう。」
さっつーのこの言葉を胸に、サメじろうはパパと仲直りできるのか!?

サメじろうが選んだのは、お別れか、それとも仲直りか。赦しあう心、家族の大切さを再確認できる、手描きのアナログイラストによる温かなアニメーション風漫画です。ほのぼのとした癒しと感動が広がる第19話を、ぜひご覧ください。


イラスト・原作:ソラガスキ

Physical AI (E)