WO2025191449 - SYSTEMS AND METHODS FOR MACHINE LEARNING-BASED GENOME ANNOTATION

National phase entry is expected:
Publication Number WO/2025/191449
Publication Date 18.09.2025
International Application No. PCT/IB2025/052553
International Filing Date 10.03.2025
Title **
[English] SYSTEMS AND METHODS FOR MACHINE LEARNING-BASED GENOME ANNOTATION
[French] SYSTÈMES ET PROCÉDÉS D'ANNOTATION DE GÉNOME BASÉE SUR UN APPRENTISSAGE AUTOMATIQUE
Applicants **
INSTADEEP LTD
BIONTECH SE
Inventors
PIERROT, Thomas
DE ALMEIDA, Bernardo P.
RICHARD, Guillaume
DALLA-TORRE, Hugo
LATERRE, Alexandre
BEGUIR, Karim
HEXEMER, Lorenz Johann Leopold
LAURENT, Stephen Jean Yvon
LANG, Maren
PANDEY, Priyanka
SAHIN, Ugur
Priority Data
63/563,903   11.03.2024   US
63/683,682   15.08.2024   US
63/701,114   30.09.2024   US
Application details
Total Number of Claims/PCT *
Number of Independent Claims *
Number of Priorities *
Number of Multi-Dependent Claims *
Number of Drawings *
Pages for Publication *
Number of Pages with Drawings *
Pages of Specification *
*
Number of Office Actions *
*
International Searching Authority
*
Recordal of a Change of the Applicant's Name/Address
*
Type of Assignment
*
Applicant's Legal Status
*
*
*
*
*
*
Entry into National Phase under
*
Patent Delivery
*
Translation

* The data is based on automatic recognition. Please verify and amend if necessary.

** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.

Quotation for National Phase entry

Country StagesTotal
China Filing, Examination, Granting5110
EPO Filing, Examination, Granting68275
Japan Filing, Examination, Granting4157
South Korea Filing, Examination, Granting6222
USA Filing, Examination, Granting23040
MasterCard Visa
Total: 106,804
Contact Us
Abstract[English] The present disclosure, among other things, provides machine-learning technologies for identifying and localizing particular genomic elements (e.g., gene elements and/or regulatory elements) within nucleotide sequences, such as DNA and/or RNA sequences. In certain embodiments, similar to the manner in which image processing methods can be used to localize particular objects in images at pixel level resolution, referred to as "segmentation," systems and methods of the present disclosure predict presence and locations of certain genomic elements within nucleotide sequences, thereby "segmenting" nucleotide sequences. Accordingly, genomic element segmentation technologies described herein may be used to generate annotations that identify and label portions of nucleotide sequences according to their predicted (e.g., via machine learning models described herein) function – e.g., as protein-coding genes, untranslated regions, splice sites, promotors, enhancers, etc. Among other things, these genomic annotations may be used to inform underlying biological processes driving diseases and facilitate development of new therapies.[French] La présente divulgation concerne, entre autres, des technologies d'apprentissage automatique pour identifier et localiser des éléments génomiques particuliers (par exemple, des éléments géniques et/ou des éléments régulateurs) dans des séquences nucléotidiques, telles que des séquences d'ADN et/ou d'ARN. Dans certains modes de réalisation, d'une manière similaire à celle selon laquelle des procédés de traitement d'image peuvent être utilisés pour localiser des objets particuliers dans des images à une résolution au niveau du pixel, appelée "segmentation", les systèmes et procédés de la présente divulgation prédisent la présence et les emplacements de certains éléments génomiques dans des séquences nucléotidiques, pour "segmenter" ainsi des séquences nucléotidiques. En conséquence, les technologies de segmentation d'éléments génomiques décrites ici peuvent être utilisées pour générer des annotations qui identifient et marquent des parties de séquences nucléotidiques selon leur fonction prédite (par exemple, par l'intermédiaire de modèles d'apprentissage automatique décrits ici), par exemple, en tant que gènes codant pour des protéines, régions non traduites, sites d'épissage, promoteurs, amplificateurs, etc. Entre autres, ces annotations génomiques peuvent être utilisées pour déterminer des processus biologiques sous-jacents entraînant des maladies, et faciliter le développement de nouvelles thérapies.