WO2025243140 - UTILIZING A LARGE LANGUAGE MODEL (LLM) TO AUTOMATICALLY CONSTRUCT A MACHINE LEARNING (ML) CLASSIFICATION MODEL
National phase entry is expected:
Publication Number
WO/2025/243140
Publication Date
27.11.2025
International Application No.
PCT/IB2025/054949
International Filing Date
12.05.2025
Title **
[English]
UTILIZING A LARGE LANGUAGE MODEL (LLM) TO AUTOMATICALLY CONSTRUCT A MACHINE LEARNING (ML) CLASSIFICATION MODEL
[French]
UTILISATION D'UN GRAND MODÈLE DE LANGAGE (LLM) POUR CONSTRUIRE AUTOMATIQUEMENT UN MODÈLE DE CLASSIFICATION D'APPRENTISSAGE MACHINE (ML)
Applicants **
VARONIS SYSTEMS, INC.
Inventors
OSI, Amit
SNEH, Ron
Priority Data
18/670,752
22.05.2024
US
Application details
| Total Number of Claims/PCT | * |
| Number of Independent Claims | * |
| Number of Priorities | * |
| Number of Multi-Dependent Claims | * |
| Number of Drawings | * |
| Pages for Publication | * |
| Number of Pages with Drawings | * |
| Pages of Specification | * |
| * | |
| Number of Office Actions | * |
| * | |
International Searching Authority |
USPTO
* |
| * | |
| Recordal of a Change of the Applicant's Name/Address |
Change of Applicant's Name and Address
* |
| Type of Assignment |
The Standard Agent's Assignment
* |
| Applicant's Legal Status |
Legal Entity
* |
| * | |
| * | |
| * | |
| * | |
| * | |
| Entry into National Phase under |
Chapter I
* |
| Patent Delivery |
Send the Letters Patent by Courier
* |
| Translation |
|
* The data is based on automatic recognition. Please verify and amend if necessary.
** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.
Quotation for National Phase entry
| Country | Stages | Total | |
|---|---|---|---|
| China | Filing, Examination, Granting | 2328 | |
| EPO | Filing, Examination, Granting | 14428 | |
| Japan | Filing, Examination, Granting | 2325 | |
| South Korea | Filing, Examination, Granting | 2454 | |
| USA | Filing, Examination, Granting | 4310 |

Total:
25,845
Contact Us
Abstract[English]
A computerized method includes: obtaining a first dataset of pre-labeled textual items, wherein each pre-labeled textual item is associated with a pre-label; feeding each of the pre-labeled textual items into a Large Language Model (LLM), and prompting it to generate textual reasoning that supports the pre-label of each pre-labeled textual item; collating the generated textual reasonings, and generating therefrom a textual instruction prompt; obtaining a second dataset of not-yet- labeled textual items; feeding each of the not-yet-labeled textual items into the LLM, and commanding it to utilize the textual instruction prompt and to generate a textual label for each of the not-yet-labeled textual items; collecting those textual items, that were labeled by the LLM, into a third dataset of LLM-labeled textual items; automatically training a Machine Language (ML) classification model on that third dataset of LLM-labeled textual items; deploying that ML classification model in a platform for classification of textual items.[French]
Un procédé informatisé consiste à : obtenir un premier ensemble de données d'éléments textuels pré-étiquetés, chaque élément textuel pré-étiqueté étant associé à une pré-étiquette ; introduire chacun des éléments textuels pré-étiquetés dans un grand modèle de langage (LLM), et lui donner l'instruction de générer un raisonnement textuel qui prend en charge la pré-étiquette de chaque élément textuel pré-étiqueté ; collationner les raisonnements textuels générés, et générer à partir de ceux-ci une instruction générative textuelle ; obtenir un deuxième ensemble de données d'éléments textuels pas encore étiquetés ; introduire chacun des éléments textuels pas encore étiquetés dans le LLM, et commander celui-ci pour qu'il utilise l'instruction générative textuelle et qu'il génère une étiquette textuelle pour chacun des éléments textuels pas encore étiquetés ; collecter ces éléments textuels, qui ont été étiquetés par le LLM, dans un troisième ensemble de données d'éléments textuels étiquetés par le LLM ; entraîner automatiquement un modèle de classification de langage machine (ML) sur ce troisième ensemble de données d'éléments textuels étiquetés par le LLM ; déployer ce modèle de classification de ML dans une plateforme pour la classification d'éléments textuels.