WO2025125288 - PROCESSING HETEROGENEOUS CONTENT IN LANGUAGE MODELS
National phase entry is expected:
Publication Number
WO/2025/125288
Publication Date
19.06.2025
International Application No.
PCT/EP2024/085604
International Filing Date
11.12.2024
Title **
[English]
PROCESSING HETEROGENEOUS CONTENT IN LANGUAGE MODELS
[French]
TRAITEMENT DE CONTENU HÉTÉROGÈNE DANS DES MODÈLES DE LANGAGE
Applicants **
IP MIND LTD
Inventors
GILL, Sharaz
NAIDU, Deepal
Priority Data
2318920.2
12.12.2023
GB
2319323.8
15.12.2023
GB
2408498.0
13.06.2024
GB
Application details
| Total Number of Claims/PCT | * |
| Number of Independent Claims | * |
| Number of Priorities | * |
| Number of Multi-Dependent Claims | * |
| Number of Drawings | * |
| Pages for Publication | * |
| Number of Pages with Drawings | * |
| Pages of Specification | * |
| * | |
| Number of Office Actions | * |
| * | |
International Searching Authority |
EPO
* |
| Recordal of a Change of the Applicant's Name/Address |
Change of Applicant's Name and Address
* |
| Type of Assignment |
The Standard Agent's Assignment
* |
| Applicant's Legal Status |
Legal Entity
* |
| * | |
| * | |
| * | |
| * | |
| * | |
| Entry into National Phase under |
Chapter I
* |
| Patent Delivery |
Send the Letters Patent by Courier
* |
| 译文 |
|
* The data is based on automatic recognition. Please verify and amend if necessary.
** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.
Quotation for National Phase entry
| Country | Stages | Total | |
|---|---|---|---|
| China | Filing, Examination, Granting | 2394 | |
| EPO | Filing, Examination, Granting | 12705 | |
| Japan | Filing, Examination, Granting | 2431 | |
| South Korea | Filing, Examination, Granting | 2703 | |
| USA | Filing, Examination, Granting | 6140 |

Total:
26,373
The term for entry into the National Phase has expired. This quotation is for informational purposes only
Contact Us
Abstract[English]
A method and system for enhancing information retrieval and understanding in language models when processing heterogeneous documents containing text, tables, figures, formulas, and images is described. The method involves traversing a document to identify specific types of content and using a language model to generate text descriptions of the identified content, including detailed interpretations. These descriptions are stored in a storage means, indexed by content type. During retrieval, the language model references these descriptions to enhance its understanding and analysis of the document, improving accuracy and contextual relevance. The system comprises a document traversing module, a language model, a storage means, and an information retrieval module, with a feedback loop for continuous improvement based on user feedback. This invention addresses limitations in current retrieval-augmented generation applications by improving the interpretation of complex technical content, thereby enhancing the performance of language models.[French]
L'invention concerne un procédé et un système pour améliorer la récupération et la compréhension d'informations dans des modèles de langage lors du traitement de documents hétérogènes contenant du texte, des tableaux, des figures, des formules et des images. Le procédé consiste à traverser un document pour identifier des types de contenu spécialisés et à utiliser un modèle de langage pour générer des descriptions de texte du contenu identifié, y compris des interprétations détaillées. Ces descriptions sont stockées dans un moyen de stockage, indexé par type de contenu. Pendant la récupération, le modèle de langue référence ces descriptions pour améliorer sa compréhension et l'analyse du document, ce qui améliore la précision et la pertinence contextuelle. Le système comprend un module de traversée de document, un modèle de langage, un moyen de stockage et un module de récupération d'informations, avec une boucle de rétroaction pour une amélioration continue en fonction des retours de l'utilisateur. La présente invention aborde des limitations dans des applications de génération augmentée-récupération actuelles par amélioration de l'interprétation d'un contenu technique complexe, ce qui permet d'améliorer les performances de modèles de langage.