WO2026125027 - DEVICE AND METHOD FOR TRAINING A CONTROL POLICY FOR A ROBOT DEVICE
National phase entry is expected:
Publication Number
WO/2026/125027
Publication Date
18.06.2026
International Application No.
PCT/EP2025/084661
International Filing Date
28.11.2025
Title **
[English]
DEVICE AND METHOD FOR TRAINING A CONTROL POLICY FOR A ROBOT DEVICE
[French]
DISPOSITIF ET PROCÉDÉ D’APPRENTISSAGE D’UNE POLITIQUE DE COMMANDE POUR UN DISPOSITIF ROBOTIQUE
Applicants **
ROBERT BOSCH GMBH
Inventors
DING, Haoran
ROZO, Leonel
Priority Data
102024211812.5
11.12.2024
DE
Application details
| Total Number of Claims/PCT | * |
| Number of Independent Claims | * |
| Number of Priorities | * |
| Number of Multi-Dependent Claims | * |
| Number of Drawings | * |
| Pages for Publication | * |
| Number of Pages with Drawings | * |
| Pages of Specification | * |
| * | |
| Number of Office Actions | * |
| * | |
International Searching Authority |
EPO
* |
| Recordal of a Change of the Applicant's Name/Address |
Change of Applicant's Name and Address
* |
| Type of Assignment |
The Standard Agent's Assignment
* |
| Applicant's Legal Status |
Legal Entity
* |
| * | |
| * | |
| * | |
| * | |
| * | |
| Entry into National Phase under |
Chapter I
* |
| Patent Delivery |
Send the Letters Patent by Courier
* |
| Translation |
|
* The data is based on automatic recognition. Please verify and amend if necessary.
** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.
Quotation for National Phase entry
| Country | Stages | Total | |
|---|---|---|---|
| China | Filing, Examination, Granting | 1880 | |
| EPO | Filing, Examination, Granting | 8113 | |
| Japan | Filing, Examination, Granting | 1968 | |
| South Korea | Filing, Examination, Granting | 1693 | |
| USA | Filing, Examination, Granting | 5340 |

Total:
18,994
Contact Us
Abstract[English]
According to various embodiments, a method for training a control policy for a robot device is described, comprising providing demonstrations, each demonstration indicating, for one or more observations, a sequence of actions of the robot device indicating how the robot device should react to the one or more observations and training a neural network to imitate, for each demonstration of the demonstrations, a reference vector field, wherein the reference vector field is the negative gradient of an energy function and following the reference vector field maps points representing action sequences to a point representing the action sequence indicated by the demonstration.[French]
Selon divers modes de réalisation, un procédé d’apprentissage d’une politique de commande pour un dispositif robotique est décrit, comprenant la fourniture de démonstrations, chaque démonstration indiquant, pour une ou plusieurs observations, une séquence d’actions du dispositif robotique indiquant comment ledit dispositif robotique doit réagir auxdites observations, et l’apprentissage d’un réseau neuronal destiné à imiter, pour chaque démonstration parmi lesdites démonstrations, un champ de vecteurs de référence, ledit champ de vecteurs de référence étant le gradient négatif d’une fonction d’énergie, le suivi dudit champ de vecteurs de référence mappant des points représentant des séquences d’actions vers un point représentant la séquence d’actions indiquée par la démonstration.