WO2023084348 - EMOTION RECOGNITION IN MULTIMEDIA VIDEOS USING MULTI-MODAL FUSION-BASED DEEP NEURAL NETWORK

National phase entry:
Publication Number WO/2023/084348
Publication Date 19.05.2023
International Application No. PCT/IB2022/060334
International Filing Date 27.10.2022
Title **
[English] EMOTION RECOGNITION IN MULTIMEDIA VIDEOS USING MULTI-MODAL FUSION-BASED DEEP NEURAL NETWORK
[French] RECONNAISSANCE D'ÉMOTION DANS DES VIDÉOS MULTIMÉDIAS À L'AIDE D'UN RÉSEAU NEURONAL PROFOND BASÉ SUR LA FUSION MULTIMODALE
Applicants **
SONY GROUP CORPORATION
Inventors
WASNIK, Pankaj
ONOE, Naoyuki
CHUDASAMA, Vishal
Priority Data
63/263,961   12.11.2021   US
17/941,787   09.09.2022   US
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
Translation

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting2490
EPO Filing, Examination, Granting11681
Japan Filing, Examination, Granting2307
South Korea Filing, Examination, Granting2418
USA Filing, Examination, Granting4740
MasterCard Visa
Total: 23,636

The term for entry into the National Phase has expired. This quotation is for informational purposes only

Contact Us
Abstract[English] A system and method of landmark detection using emotion recognition in multimedia videos using multi-modal fusion based deep neural network is provided. The system includes circuitry and a memory configured to store a multimodal fusion network which includes one or more feature extractors, a network of transformer encoders, a fusion attention network, and an output network coupled to the fusion attention network. The system inputs a multimodal input to the one or more feature extractors. The multimodal input is associated with an utterance depicted in one or more videos. The system generates input embeddings as an output of the one or more feature extractors for the input and further generates a set of emotion-relevant features based on the input embeddings. The system further generates a fused-feature representation of the set of emotion-relevant features and predicts an emotion label for the utterance based on fused-feature representation.[French] L'invention concerne un système et un procédé de détection de points de repère à l'aide d'une reconnaissance d'émotion dans des vidéos multimédias à l'aide d'un réseau neuronal profond basé sur la fusion multimodale. Le système comprend un ensemble de circuits et une mémoire configurée pour stocker un réseau de fusion multimodal qui comprend un ou plusieurs extracteurs de caractéristiques, un réseau de codeurs de transformateur, un réseau d'attention de fusion et un réseau de sortie couplé au réseau d'attention de fusion. Le système entre une entrée multimodale dans le ou les extracteurs de caractéristiques. L'entrée multimodale est associée à un énoncé représenté dans une ou plusieurs vidéos. Le système génère des intégrations d'entrée en tant que sortie du ou des extracteurs de caractéristiques pour l'entrée et génère en outre un ensemble de caractéristiques appropriées à une émotion sur la base des intégrations d'entrée. Le système génère en outre une représentation de caractéristiques fusionnées de l'ensemble de caractéristiques appropriées à une émotion et prédit une étiquette d'émotion pour l'énoncé sur la base de la représentation de caractéristiques fusionnées.

Rejoining the server...