WO2026109928 - SYSTEM AND METHOD FOR EXTRACTING AND CATEGORIZING INFORMATION FROM ONLINE SOURCES

National phase entry is expected:
Publication Number WO/2026/109928
Publication Date 28.05.2026
International Application No. PCT/IB2024/061780
International Filing Date 25.11.2024
Title **
[English] SYSTEM AND METHOD FOR EXTRACTING AND CATEGORIZING INFORMATION FROM ONLINE SOURCES
[French] SYSTÈME ET PROCÉDÉ D'EXTRACTION ET DE CATÉGORISATION D'INFORMATIONS À PARTIR DE SOURCES EN LIGNE
Applicants **
6SENSE INSIGHTS INC.
Inventors
SELVARAJ, Ernest Kirubakaran
GOLSEFID, Samira
BAJARIA, Viral Tarun
CHILLOJI, Satish Arjun
SHAH, Akshay Rajendra
SEKAR, Amresh
SUNWALKA, Shubham Kumar
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
Translation

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting1966
EPO Filing, Examination, Granting11797
Japan Filing, Examination, Granting1998
South Korea Filing, Examination, Granting1711
USA Filing, Examination, Granting4310
MasterCard Visa
Total: 21,782
Contact Us
Abstract[English] A system and method for efficiently extracting and categorizing business information from online sources is disclosed. The system comprises a web crawler that obtains company domains from a database and collects depth-1 URLs from company homepages. A classification model, utilizing a fine-tuned BERT architecture, predicts which URLs contain relevant information for generating tags. A content extractor then extracts content from these predicted URLs using one or more modules. Finally, a large language model (LLM) processes the extracted content and generates tags using custom prompts designed for each tag category. These prompts are tailored to the nature of the extracted content, enhancing the context provided to the LLM. This multi-stage approach addresses challenges in processing large-scale, unstructured business data from diverse web sources, potentially offering improved efficiency, scalability, and accuracy in automated business intelligence gathering.[French] L'invention concerne un système et un procédé d'extraction et de catégorisation efficaces d'informations commerciales à partir de sources en ligne. Le système comprend un robot d'indexation Web qui obtient des domaines d'entreprise à partir d'une base de données et collecte des URL de profondeur 1 à partir de pages d'accueil de sociétés. Un modèle de classification, utilisant une architecture BERT affinée, prédit quelles URL contiennent des informations pertinentes pour générer des étiquettes. Un extracteur de contenu extrait ensuite du contenu de ces URL prédites à l'aide d'un ou de plusieurs modules. Enfin, un grand modèle de langage (LLM) traite le contenu extrait et génère des étiquettes à l'aide d'invites personnalisées conçues pour chaque catégorie d'étiquette. Ces invites sont adaptées à la nature du contenu extrait, améliorant le contexte fourni au LLM. Cette approche en plusieurs étapes résout des défis liés au traitement de données commerciales non structurées à grande échelle provenant de diverses sources Web, offrant potentiellement une efficacité, une extensibilité et une précision améliorées dans la collecte automatisée de renseignements commerciaux.

Rejoining the server...