WO2025125288 - PROCESSING HETEROGENEOUS CONTENT IN LANGUAGE MODELS

National phase entry is expected:
Publication Number WO/2025/125288
Publication Date 19.06.2025
International Application No. PCT/EP2024/085604
International Filing Date 11.12.2024
Title **
[English] PROCESSING HETEROGENEOUS CONTENT IN LANGUAGE MODELS
[French] TRAITEMENT DE CONTENU HÉTÉROGÈNE DANS DES MODÈLES DE LANGAGE
Applicants **
IP MIND LTD
Inventors
GILL, Sharaz
NAIDU, Deepal
Priority Data
2318920.2   12.12.2023   GB
2319323.8   15.12.2023   GB
2408498.0   13.06.2024   GB
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
译文

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting2394
EPO Filing, Examination, Granting12705
Japan Filing, Examination, Granting2431
South Korea Filing, Examination, Granting2703
USA Filing, Examination, Granting6140
MasterCard Visa
Total: 26,373

The term for entry into the National Phase has expired. This quotation is for informational purposes only

Contact Us
Abstract[English] A method and system for enhancing information retrieval and understanding in language models when processing heterogeneous documents containing text, tables, figures, formulas, and images is described. The method involves traversing a document to identify specific types of content and using a language model to generate text descriptions of the identified content, including detailed interpretations. These descriptions are stored in a storage means, indexed by content type. During retrieval, the language model references these descriptions to enhance its understanding and analysis of the document, improving accuracy and contextual relevance. The system comprises a document traversing module, a language model, a storage means, and an information retrieval module, with a feedback loop for continuous improvement based on user feedback. This invention addresses limitations in current retrieval-augmented generation applications by improving the interpretation of complex technical content, thereby enhancing the performance of language models.[French] L'invention concerne un procédé et un système pour améliorer la récupération et la compréhension d'informations dans des modèles de langage lors du traitement de documents hétérogènes contenant du texte, des tableaux, des figures, des formules et des images. Le procédé consiste à traverser un document pour identifier des types de contenu spécialisés et à utiliser un modèle de langage pour générer des descriptions de texte du contenu identifié, y compris des interprétations détaillées. Ces descriptions sont stockées dans un moyen de stockage, indexé par type de contenu. Pendant la récupération, le modèle de langue référence ces descriptions pour améliorer sa compréhension et l'analyse du document, ce qui améliore la précision et la pertinence contextuelle. Le système comprend un module de traversée de document, un modèle de langage, un moyen de stockage et un module de récupération d'informations, avec une boucle de rétroaction pour une amélioration continue en fonction des retours de l'utilisateur. La présente invention aborde des limitations dans des applications de génération augmentée-récupération actuelles par amélioration de l'interprétation d'un contenu technique complexe, ce qui permet d'améliorer les performances de modèles de langage.