WO2025253261 - ZERO-SHOT OPEN-VOCABULARY 3D AUTO-LABELING USING VISUAL FOUNDATION MODELS

National phase entry is expected:
Publication Number WO/2025/253261
Publication Date 11.12.2025
International Application No. PCT/IB2025/055624
International Filing Date 31.05.2025
Title **
[English] ZERO-SHOT OPEN-VOCABULARY 3D AUTO-LABELING USING VISUAL FOUNDATION MODELS
[French] ÉTIQUETAGE AUTOMATIQUE 3D À VOCABULAIRE OUVERT À ZÉRO COUP À L'AIDE DE MODÈLES DE FONDATION VISUELLE
Applicants **
ROBERT BOSCH GMBH
Inventors
ZHAO, Cheng
WANG, Ruoyu
GUO, Yuliang
HUANG, Xinyu
REN, Liu
Priority Data
18/731,578   03.06.2024   US
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
译文

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting2250
EPO Filing, Examination, Granting11793
Japan Filing, Examination, Granting2371
South Korea Filing, Examination, Granting2511
USA Filing, Examination, Granting5140
MasterCard Visa
Total: 24,065
Contact Us
Abstract[English] Zero-shot open-vocabulary 3D auto-labeling is performed using visual foundation models (VFMs). Multi-view 2D images of an environment and corresponding 3D LiDAR points of the environment are received. 2D semantic knowledge is extracted from the multi-view 2D images in close-set and open-set detection branches. 3D spatial-temporal prompts are generated via clustering and tracking of the 3D LiDAR points. The 3D spatial-temporal prompts and the 2D semantic knowledge are used for mapping the 2D semantic knowledge to a plurality of clusters of the 3D LiDAR points, thereby producing labeled 3D LiDAR points defining a 3D semantic segmentation of the 3D LiDAR points. One or more downstream applications are performed using the labeled 3D LiDAR points.[French] Un étiquetage automatique 3D à vocabulaire ouvert à zéro coup est réalisé à l'aide de modèles de fondation visuelle (VFM). Des images 2D multi-vues d'un environnement et des points LiDAR 3D correspondants de l'environnement sont reçus. Des connaissances sémantiques 2D sont extraites des images 2D multi-vues dans des branches de détection d'ensemble fermé et d’ensemble ouvert. Des invites spatio-temporelles 3D sont générées par regroupement et suivi des points LiDAR 3D. Les invites spatio-temporelles 3D et les connaissances sémantiques 2D sont utilisées pour mapper les connaissances sémantiques 2D à une pluralité de groupes des points LiDAR 3D, produisant ainsi des points LiDAR 3D marqués définissant une segmentation sémantique 3D des points LiDAR 3D. Une ou plusieurs applications en aval sont effectuées à l'aide des points LiDAR 3D marqués.