WO2025248322 - AUDIO PROCESSING METHOD, MODEL TRAINING METHOD AND APPARATUSES
National phase entry is expected:
Publication Number
WO/2025/248322
Publication Date
04.12.2025
International Application No.
PCT/IB2025/052709
International Filing Date
14.03.2025
Title **
[English]
AUDIO PROCESSING METHOD, MODEL TRAINING METHOD AND APPARATUSES
[French]
PROCÉDÉ DE TRAITEMENT AUDIO, PROCÉDÉ ET APPAREILS D'ENTRAÎNEMENT DE MODÈLE
Applicants **
ALIBABA INNOVATION PRIVATE LIMITED
Inventors
YE, Jiaqi
ZHAO, Shengkui
HUANG, Dianwen
CHNG, Eng Siong
MA, Bin
Priority Data
10202401512S
28.05.2024
SG
Application details
| Total Number of Claims/PCT | * |
| Number of Independent Claims | * |
| Number of Priorities | * |
| Number of Multi-Dependent Claims | * |
| Number of Drawings | * |
| Pages for Publication | * |
| Number of Pages with Drawings | * |
| Pages of Specification | * |
| * | |
| Number of Office Actions | * |
| * | |
International Searching Authority |
IPOS
* |
| Recordal of a Change of the Applicant's Name/Address |
Change of Applicant's Name and Address
* |
| Type of Assignment |
The Standard Agent's Assignment
* |
| Applicant's Legal Status |
Legal Entity
* |
| * | |
| * | |
| * | |
| * | |
| * | |
| Entry into National Phase under |
Chapter I
* |
| Patent Delivery |
Send the Letters Patent by Courier
* |
| Translation |
|
* The data is based on automatic recognition. Please verify and amend if necessary.
** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.
Quotation for National Phase entry
| Country | Stages | Total | |
|---|---|---|---|
| China | Filing, Examination, Granting | 2435 | |
| EPO | Filing, Examination, Granting | 14820 | |
| Japan | Filing, Examination, Granting | 2309 | |
| South Korea | Filing, Examination, Granting | 2469 | |
| USA | Filing, Examination, Granting | 7140 |

Total:
29,173
Contact Us
Abstract[English]
The present disclosure provides audio processing methods, model training methods and apparatuses, devices, and a storage medium. The method includes: inputting mixed audio data into an encoder of a pre-trained neural audio codec (NAC) to obtain a first embedding of the mixed audio data, where the mixed audio data is a mixture of multiple pieces of audio data from multiple acoustic sources, and the first embedding is a representation of an audio feature of the mixture; where the pre-trained NAC is trained with a set of sample mixed audio data for audio separation; performing audio separation on the first embedding with a pre-trained separator to obtain at least one second embedding, where the at least one second embedding corresponds to at least one of the multiple pieces of audio data one-by-one, each of the at least one second embedding is a representation of an audio feature of corresponding audio data, and the at least one second embedding is used for reconstruction of the at least one of the multiple pieces of audio data with a decoder of the pre-trained NAC.[French]
La présente divulgation concerne des procédés de traitement audio, des procédés et des appareils d’entraînement de modèle, des dispositifs et un support de stockage. Le procédé consiste à : entrer des données audio mélangées dans un codeur d'un codec audio neuronal (NAC) pré-entraîné pour obtenir une première incorporation des données audio mélangées, les données audio mélangées étant un mélange de multiples éléments de données audio provenant de multiples sources acoustiques, et la première incorporation étant une représentation d'une caractéristique audio du mélange, le NAC pré-entraîné étant entraîné sur un ensemble de données audio mélangées d'échantillon à des fins de séparation audio ; réaliser une séparation audio sur la première incorporation au moyen d’un séparateur pré-entraîné pour obtenir au moins une seconde incorporation, l'au moins une seconde incorporation correspondant à au moins l'un des multiples éléments de données audio sur une base individuelle, chacune de l'au moins une seconde incorporation étant une représentation d'une caractéristique audio de données audio correspondantes, et l'au moins une seconde incorporation étant utilisée pour la reconstruction de l'au moins un des multiples éléments de données audio au moyen d’un décodeur du NAC pré-entraîné.