WO2025243140 - UTILIZING A LARGE LANGUAGE MODEL (LLM) TO AUTOMATICALLY CONSTRUCT A MACHINE LEARNING (ML) CLASSIFICATION MODEL

National phase entry is expected:
Publication Number WO/2025/243140
Publication Date 27.11.2025
International Application No. PCT/IB2025/054949
International Filing Date 12.05.2025
Title **
[English] UTILIZING A LARGE LANGUAGE MODEL (LLM) TO AUTOMATICALLY CONSTRUCT A MACHINE LEARNING (ML) CLASSIFICATION MODEL
[French] UTILISATION D'UN GRAND MODÈLE DE LANGAGE (LLM) POUR CONSTRUIRE AUTOMATIQUEMENT UN MODÈLE DE CLASSIFICATION D'APPRENTISSAGE MACHINE (ML)
Applicants **
VARONIS SYSTEMS, INC.
Inventors
OSI, Amit
SNEH, Ron
Priority Data
18/670,752   22.05.2024   US
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
Translation

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting2328
EPO Filing, Examination, Granting14428
Japan Filing, Examination, Granting2325
South Korea Filing, Examination, Granting2454
USA Filing, Examination, Granting4310
MasterCard Visa
Total: 25,845
Contact Us
Abstract[English] A computerized method includes: obtaining a first dataset of pre-labeled textual items, wherein each pre-labeled textual item is associated with a pre-label; feeding each of the pre-labeled textual items into a Large Language Model (LLM), and prompting it to generate textual reasoning that supports the pre-label of each pre-labeled textual item; collating the generated textual reasonings, and generating therefrom a textual instruction prompt; obtaining a second dataset of not-yet- labeled textual items; feeding each of the not-yet-labeled textual items into the LLM, and commanding it to utilize the textual instruction prompt and to generate a textual label for each of the not-yet-labeled textual items; collecting those textual items, that were labeled by the LLM, into a third dataset of LLM-labeled textual items; automatically training a Machine Language (ML) classification model on that third dataset of LLM-labeled textual items; deploying that ML classification model in a platform for classification of textual items.[French] Un procédé informatisé consiste à : obtenir un premier ensemble de données d'éléments textuels pré-étiquetés, chaque élément textuel pré-étiqueté étant associé à une pré-étiquette ; introduire chacun des éléments textuels pré-étiquetés dans un grand modèle de langage (LLM), et lui donner l'instruction de générer un raisonnement textuel qui prend en charge la pré-étiquette de chaque élément textuel pré-étiqueté ; collationner les raisonnements textuels générés, et générer à partir de ceux-ci une instruction générative textuelle ; obtenir un deuxième ensemble de données d'éléments textuels pas encore étiquetés ; introduire chacun des éléments textuels pas encore étiquetés dans le LLM, et commander celui-ci pour qu'il utilise l'instruction générative textuelle et qu'il génère une étiquette textuelle pour chacun des éléments textuels pas encore étiquetés ; collecter ces éléments textuels, qui ont été étiquetés par le LLM, dans un troisième ensemble de données d'éléments textuels étiquetés par le LLM ; entraîner automatiquement un modèle de classification de langage machine (ML) sur ce troisième ensemble de données d'éléments textuels étiquetés par le LLM ; déployer ce modèle de classification de ML dans une plateforme pour la classification d'éléments textuels.

Rejoining the server...