WO2025219798 - UTILIZING A LARGE LANGUAGE MODEL (LLM) FOR LABELING DATA-ITEMS FOR TRAINING A MACHINE LEARNING (ML) MODEL
National phase entry is expected:
Publication Number
WO/2025/219798
Publication Date
23.10.2025
International Application No.
PCT/IB2025/053532
International Filing Date
03.04.2025
Title **
[English]
UTILIZING A LARGE LANGUAGE MODEL (LLM) FOR LABELING DATA-ITEMS FOR TRAINING A MACHINE LEARNING (ML) MODEL
[French]
UTILISATION D'UN GRAND MODÈLE DE LANGAGE (LLM) POUR ÉTIQUETER DES ÉLÉMENTS DE DONNÉES POUR ENTRAÎNER UN MODÈLE D'APPRENTISSAGE AUTOMATIQUE (ML)
Applicants **
VARONIS SYSTEMS, INC.
Inventors
OSI, Amit
SNEH, Ron
Priority Data
18/636,297
16.04.2024
US
Application details
| Total Number of Claims/PCT | * |
| Number of Independent Claims | * |
| Number of Priorities | * |
| Number of Multi-Dependent Claims | * |
| Number of Drawings | * |
| Pages for Publication | * |
| Number of Pages with Drawings | * |
| Pages of Specification | * |
| * | |
| Number of Office Actions | * |
| * | |
International Searching Authority |
USPTO
* |
| * | |
| Recordal of a Change of the Applicant's Name/Address |
Change of Applicant's Name and Address
* |
| Type of Assignment |
The Standard Agent's Assignment
* |
| Applicant's Legal Status |
Legal Entity
* |
| * | |
| * | |
| * | |
| * | |
| * | |
| Entry into National Phase under |
Chapter I
* |
| Patent Delivery |
Send the Letters Patent by Courier
* |
| Translation |
|
* The data is based on automatic recognition. Please verify and amend if necessary.
** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.
Quotation for National Phase entry
| Country | Stages | Total | |
|---|---|---|---|
| China | Filing, Examination, Granting | 2282 | |
| EPO | Filing, Examination, Granting | 14233 | |
| Japan | Filing, Examination, Granting | 2325 | |
| South Korea | Filing, Examination, Granting | 2441 | |
| USA | Filing, Examination, Granting | 4910 |

Total:
26,191
Contact Us
Abstract[English]
A Large Language Model (LLM) is configured to automatically label non-labeled textual data- items for the purpose of creating a training dataset for training a Machine Learning (ML) model. The ML model is thus trained on LLM-labeled textual data-items; and the ML model can be deployed to classify new or incoming documents or messages or other textual data- items. Additionally, a Vision and Language Model (VLM) or a Large Multimodal Model (LMM) or a large multiple-modalities model (LMM) can process data from two or more modalities (visual data, textual data), and is configured to automatically label non-labeled images for the purpose of creating a training dataset for training a Deep Neural Network (DNN) or a Deep Convolutional Neural Network (Deep CNN) model. The DNN model is thus trained on VLM-labeled images; and the DNN model can be deployed to classify new or incoming images.[French]
Un grand modèle de langage (LLM) est configuré pour étiqueter automatiquement des éléments de données textuels non étiquetés dans le but de créer un ensemble de données d'entraînement pour entraîner un modèle d'apprentissage automatique (ML). Le modèle ML est ainsi entraîné sur des éléments de données textuels étiquetés par LLM ; et le modèle ML peut être déployé pour classifier des documents ou des messages nouveaux ou entrants ou d'autres éléments de données textuels. De plus, un modèle de vision et de langage (VLM), un grand modèle multimodal (LMM) ou un grand modèle à modalités multiples (LMM) peuvent traiter des données à partir de deux modalités ou plus (données visuelles, données textuelles), et sont configurés pour étiqueter automatiquement des images non étiquetées dans le but de créer un ensemble de données d'entraînement pour entraîner un réseau neuronal profond (DNN) ou un modèle de réseau neuronal convolutif profond (CNN profond). Le modèle DNN est ainsi entraîné sur des images étiquetées par VLM ; et le modèle DNN peut être déployé pour classifier de nouvelles images ou des images entrantes.