WO2025149990 - VISUAL SPEECH RECOGNITION BASED ON LIP MOVEMENTS USING GENERATIVE ARTIFICIAL INTELLIGENCE (AI) MODEL

National phase entry is expected:
Publication Number WO/2025/149990
Publication Date 17.07.2025
International Application No. PCT/IB2025/050340
International Filing Date 12.01.2025
Title **
[English] VISUAL SPEECH RECOGNITION BASED ON LIP MOVEMENTS USING GENERATIVE ARTIFICIAL INTELLIGENCE (AI) MODEL
[French] RECONNAISSANCE VISUELLE DE LA PAROLE SUR LA BASE DE MOUVEMENTS DE LÈVRES À L'AIDE D'UN MODÈLE D'INTELLIGENCE ARTIFICIELLE (IA) GÉNÉRATIVE
Applicants **
SONY GROUP CORPORATION
Inventors
LEE, Jong Hwa
COSTELA, Francisco
WISECAVER, Paul
Priority Data
63/619,871   11.01.2024   US
18/808,370   19.08.2024   US
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
Translation

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting2344
EPO Filing, Examination, Granting11181
Japan Filing, Examination, Granting2327
South Korea Filing, Examination, Granting2400
USA Filing, Examination, Granting4740
MasterCard Visa
Total: 22,992
Contact Us
Abstract[English] An electronic device and a method for implementation for visual speech recognition based on lip movements. The electronic device receives a set of images including one or more human speaker and applies a first machine learning (ML) model on the received set of images. The electronic device determines a first set of words spoken by the one or more human speakers based on the application of the first ML model. The determined first set of words corresponds to lip movements of the one or more human speakers. The electronic device applies a first generative Artificial Intelligence (AI) model on the determined first set of words. The electronic device predicts a first sentence corresponding to the determined first set of words spoken by the one or more human speakers, based on the application of the first generative AI model.[French] L'invention concerne un dispositif électronique et un procédé de mise en œuvre pour la reconnaissance visuelle de la parole sur la base de mouvements de lèvres. Le dispositif électronique reçoit un ensemble d'images comprenant un ou plusieurs locuteurs humains et applique un premier modèle d'apprentissage automatique (ML) sur l'ensemble d'images reçu. Le dispositif électronique détermine un premier ensemble de mots prononcés par le ou les locuteurs humains sur la base de l'application du premier modèle ML. Le premier ensemble de mots déterminé correspond à des mouvements de lèvres du ou des locuteurs humains. Le dispositif électronique applique un premier modèle d'intelligence artificielle (IA) générative sur le premier ensemble de mots déterminé. Le dispositif électronique prédit une première phrase correspondant au premier ensemble déterminé de mots prononcés par le ou les locuteurs humains, sur la base de l'application du premier modèle d'IA générative.

Rejoining the server...