WO2026103193 - SYSTEMS AND METHODS FOR SERVICE LEVEL AGREEMENTS FOR FOUNDATION MODEL APPLICATIONS
National phase entry is expected:
Publication Number
WO/2026/103193
Publication Date
21.05.2026
International Application No.
PCT/CN2025/108150
International Filing Date
11.07.2025
Title **
[English]
SYSTEMS AND METHODS FOR SERVICE LEVEL AGREEMENTS FOR FOUNDATION MODEL APPLICATIONS
[French]
SYSTÈMES ET PROCÉDÉS POUR CONTRATS DE NIVEAU DE SERVICE POUR APPLICATIONS DE MODÈLES DE BASE
Applicants **
HUAWEI TECHNOLOGIES CO., LTD.
Inventors
THANGARAJAH, Kishanthan
LEUNG, Arthur
ZHANG, Haoxiang
CHEN, Boyuan
HASSAN, Ahmed E.
Priority Data
63/719,412
12.11.2024
US
19/197,315
02.05.2025
US
Application details
| Total Number of Claims/PCT | * |
| Number of Independent Claims | * |
| Number of Priorities | * |
| Number of Multi-Dependent Claims | * |
| Number of Drawings | * |
| Pages for Publication | * |
| Number of Pages with Drawings | * |
| Pages of Specification | * |
| * | |
| Number of Office Actions | * |
| * | |
International Searching Authority |
CNIPA
* |
| Recordal of a Change of the Applicant's Name/Address |
Change of Applicant's Name and Address
* |
| Type of Assignment |
The Standard Agent's Assignment
* |
| Applicant's Legal Status |
Legal Entity
* |
| * | |
| * | |
| * | |
| * | |
| * | |
| Entry into National Phase under |
Chapter I
* |
| Patent Delivery |
Send the Letters Patent by Courier
* |
| 译文 |
|
* The data is based on automatic recognition. Please verify and amend if necessary.
** IP-Coster compiles data from publicly available sources. If this data includes your personal information, you can contact us to request its removal.
Quotation for National Phase entry
| Country | Stages | Total | |
|---|---|---|---|
| China | Filing, Examination, Granting | 2042 | |
| EPO | Filing, Examination, Granting | 14527 | |
| Japan | Filing, Examination, Granting | 2307 | |
| South Korea | Filing, Examination, Granting | 2502 | |
| USA | Filing, Examination, Granting | 4740 |

Total:
26,118
Contact Us
Abstract[English]
Systems and methods are described for scheduling and/or resource provisioning for foundation model applications. The scheduler may involve determining a slack from a performance target and an amount of consumed resources for a workflow request; determining an available resource for at least one replica; selecting a replica, from the at least one replica, associated with a maximum amount of the available resource; and routing at least one machine learning model to a task queue associated with the selected replica for execution. The resource provisioner may involve determining a remaining slack from an associated performance target and an amount of consumed resources for a workflow request; determining a remaining resource to complete the workflow request; tracking a slack violation amount when the remaining resource exceeds the remaining slack; and increasing a number of replicas processing the workflow request based at least on the slack violation amount.[French]
L'invention concerne des systèmes et des procédés de planification et/ou de fourniture de ressources pour des applications de modèles de base. Le planificateur peut déterminer des capacités à partir d'une cible de performance et d'une quantité de ressources consommées pour une demande de flux de travail ; déterminer une ressource disponible pour au moins une réplique ; sélectionner une réplique, associée à une quantité maximale de la ressource disponible, à partir de ladite ou desdites répliques ; et acheminer au moins un modèle d'apprentissage automatique vers une file d'attente de tâches associée à la réplique sélectionnée à des fins d'exécution. Le fournisseur de ressources peut déterminer des capacités encore inutilisées à partir d'une cible de performance associée et d'une quantité de ressources consommées pour une demande de flux de travail ; déterminer une ressource restante pour finaliser la demande de flux de travail ; suivre une quantité d'atteintes aux capacités lorsque la ressource restante dépasse les capacités inutilisées ; et augmenter le nombre de répliques traitant la demande de flux de travail sur la base au moins de la quantité d'atteintes aux capacités.