Invalidity dossier

US 11650968

Systems and methods for predictive early stopping in neural network training

Current assignee: Unified Patents

Added 5/14/2026, 6:01:43 AM

At a glancePTAB challenged1 lawsuit on fileasserted by Unified PatentsHigh-Tech (T)

Active provider: Google · gemini-2.5-flash

Patent summary

Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.

✓ Generated

Here's a concise summary of US Patent 11650968:

Title: Systems and methods for predictive early stopping in neural network training

Assignee: Comet Ml Inc.

Inventors: Dhruv Nair, Gideon Mendels, Nimrod Lahav

Filing Date: 2019-12-09

Issue Date: 2023-05-16

Abstract: Systems and methods may train neural networks (NNs) and determine when to stop training to not waste computing or other resources when improvement is no longer likely. After a training period for a NN, a model trained using training data from other NNs may return a probability of improvement in the loss of the NN or a probability that the likely best loss of the NN is lower than the best loss of the other NNs for which hyperparameters have been chosen. Training may be stopped if the probability is less than a threshold, or a wait value is greater than a wait threshold.

Plain-Language Overview of Independent Claims:

The patent has several independent claims. Here's an overview of key independent claims in plain language:

  • Claim 1: This claim describes a computer-implemented method for predictively stopping the training of a neural network (NN). The method involves:

    1. Training a target NN over a series of training periods (epochs), computing a loss for each period.
    2. Using a predictive model, which was trained on loss data from other NNs (not the target NN), to determine a probability of improvement in the target NN's loss. This model takes the target NN's parameters and current/historic loss data as input.
    3. Stopping the target NN's training if this determined probability of improvement is below a certain threshold, OR if a "wait value" (a counter for non-improvement) exceeds another threshold.
  • Claim 10: This claim describes a computer system configured to perform the method outlined in Claim 1. It includes a processor, memory, and code which, when executed, carries out the steps of training the NN, determining the probability of improvement using the trained model, and stopping the training based on the thresholds.

  • Claim 16: This claim describes a computer program product (e.g., a non-transitory computer-readable medium) that stores instructions which, when executed by a processor, cause the processor to perform the method of predictively stopping NN training as described in Claim 1.

CAFC 2026 Dockets Search:

A search of CAFC 2026 dockets for patent number 11650968 did not return any specific litigation related to this patent. The search results provided general information about CAFC decisions and scheduled cases but did not mention US11650968B2.

Generated 5/16/2026, 6:47:35 PM

Cases on file (1)

Group view →

Specific litigation cases in our database that name US patent 11650968. The free-form analysis below may also discuss cases beyond this list.

Litigation summary

Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.

✓ Generated

Based on available information, there is one known litigation involving US patent 11650968.

1. PTAB Case IPR2025-01063

Additionally, the Google Patents entry for US11650968 indicates "Family has litigation" with a link to Darts-ip. While Darts-ip is a platform for global patent case data, the provided search snippets do not offer specific details about any other litigation directly involving US patent 11650968 within its patent family. Without access to the Darts-ip database, further specific case information for US11650968 from this source cannot be provided.

Generated 5/16/2026, 6:47:48 PM

Proceedings on file (1)

All PTAB activity →

AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.

Current assignee: Unified Patents

1 settled
Terminated-Settled
Filed
Jun 5, 2025
Last modified
Oct 2, 2025
Petitioner
Perplexity AI, Inc.
Inventor
Dhruv NAIR et al

PTAB challenges

AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.

✓ Generated

Proceedings overview

One AIA trial proceeding has been filed against US Patent 11650968, which was subsequently terminated due to settlement. This leaves the patent's claims untested by a final written decision, indicating that the patent has not been formally hardened or narrowed through PTAB trials.

IPR2025-01063 — Perplexity AI, Inc. v. Dhruv NAIR et al

  • Type: Inter Partes Review
  • Filed: 2025-06-05
  • Status: Terminated-Settled – The proceeding concluded due to a settlement between Perplexity AI, Inc. and Dhruv Nair et al. (representing the patent owner Comet ML Inc.), rather than a final decision on the merits.
  • Judge panel: Not publicly available without access to the full PTAB docket for this case.
  • Petition grounds: Specific claims challenged, prior art references, and statutory bases (§ 102 for anticipation / § 103 for obviousness / § 112 for indefiniteness/written description/enablement) are not yet publicly detailed in the provided information or easily discoverable without full PTAB docket access.
  • Institution decision: The outcome of the institution decision (instituted, denied, or partially instituted) and its date are not immediately available without access to the full PTAB docket for this case. Given the settlement, it is possible the case settled before an institution decision was rendered or shortly thereafter.
  • Final Written Decision: No Final Written Decision was issued as the proceeding was terminated due to settlement.
  • Settlement / termination: The proceeding was terminated due to settlement on 2025-10-02. The specific terms of the settlement are typically confidential and are not publicly disclosed.
  • Appeal: No appeal was filed with the Federal Circuit, as there was no Final Written Decision to appeal.
  • Defensive value: The settlement of IPR2025-01063 means that the validity of patent US11650968's claims was not adjudicated by the PTAB. For a defendant, this means the patent has not been "hardened" by surviving a validity challenge, but it also means no claims were invalidated. The specific reasons for the settlement are unknown, which limits the insights into the strength or weakness of the patent claims.

Strategic summary

Patent US11650968 currently has all its claims untested by a PTAB Final Written Decision. The sole IPR proceeding, IPR2025-01063, initiated by Perplexity AI, Inc., was terminated due to a settlement, leaving no claims formally canceled or sustained by the PTAB. This means that, from a PTAB perspective, the patent's claims remain as issued, without any official narrowing or confirmation of validity by the Board.

Regarding the estoppel landscape, since IPR2025-01063 was terminated due to settlement rather than a Final Written Decision, § 315(e)(2) estoppel provisions likely do not apply to the petitioner (Perplexity AI, Inc.) or its privies in the same way they would after an FWD. The exact scope of any waiver or agreement stemming from the confidential settlement is unknown, but without a merits-based decision, a new challenger would theoretically still be able to raise similar prior art grounds if they chose to file a new IPR.

There are no apparent pattern signals such as multiple IPRs from the same petitioner or aggressive appeals by the patent owner. The involvement of Perplexity AI, Inc. as a petitioner suggests a potential interest from an entity in the AI/ML space, but the settlement prevents drawing conclusions about the patent's susceptibility to challenge.

Recommended next steps

Since IPR2025-01063 was terminated by settlement, there is no Final Written Decision to link to or quote. The PTAB record on Unified Patents portal confirms the settlement.

For a defendant currently facing assertion of US11650968, the fact that the patent claims have not been adjudicated for validity by the PTAB means that any infringement theories are currently based on claims that have not been vetted through an inter partes proceeding. A potential defendant could consider initiating a new IPR if they have compelling prior art and a strong invalidity theory, as the previous settlement does not create a strong estoppel barrier against others. However, the reasons for the prior settlement should be investigated if possible, as they may offer insight into the patent owner's strategy or the perceived strength of the prior art.## Proceedings overview
One AIA trial proceeding has been filed against US Patent 11650968, which was subsequently terminated due to settlement. This leaves the patent's claims untested by a final written decision, indicating that the patent has not been formally hardened or narrowed through PTAB trials.

IPR2025-01063 — Perplexity AI, Inc. v. Dhruv NAIR et al

  • Type: Inter Partes Review
  • Filed: 2025-06-05
  • Status: Terminated-Settled – The proceeding concluded due to a settlement between Perplexity AI, Inc. and Dhruv Nair et al. (representing the patent owner Comet ML Inc.), rather than a final decision on the merits.
  • Judge panel: The specific judge panel for this proceeding is not publicly available in the provided information.
  • Petition grounds: The specific claims challenged, prior art references, and statutory bases (§ 102 / § 103 / § 112) are not publicly available in the provided information. IPR petitions typically challenge claims under 35 U.S.C. §§ 102 or 103.
  • Institution decision: The outcome of the institution decision (instituted, denied, or partially instituted) and its date are not publicly available in the provided information. Given the proceeding's "Terminated-Settled" status, it is possible the case settled before an institution decision was rendered, or shortly thereafter. The USPTO Director gained increased authority over institution decisions, including discretionary denials, effective October 20, 2025. This occurred after the IPR was filed (June 5, 2025) but before its termination (October 2, 2025).
  • Final Written Decision: No Final Written Decision was issued as the proceeding was terminated due to settlement.
  • Settlement / termination: The proceeding was terminated due to settlement on 2025-10-02. The specific terms of the settlement are typically confidential and are not publicly disclosed.
  • Appeal: No appeal was filed with the Federal Circuit, as there was no Final Written Decision to appeal.
  • Defensive value: The settlement of IPR2025-01063 means that the validity of patent US11650968's claims was not adjudicated by the PTAB. For a defendant, this implies that the patent has not been "hardened" by surviving a validity challenge, but it also means no claims were invalidated. The specific reasons for the settlement are unknown, which limits insights into the strength or weakness of the patent claims or the petitioner's prior art.

Strategic summary

Patent US11650968 currently has all its claims untested by a PTAB Final Written Decision. The sole IPR proceeding, IPR2025-01063, initiated by Perplexity AI, Inc., was terminated due to a settlement, leaving no claims formally canceled or sustained by the PTAB. This means that, from a PTAB perspective, the patent's claims remain as issued, without any official narrowing or confirmation of validity by the Board.

Regarding the estoppel landscape, since IPR2025-01063 was terminated due to settlement rather than a Final Written Decision, § 315(e)(2) estoppel provisions, which bar petitioners (and their privies) from raising any ground they raised or reasonably could have raised, likely do not apply. The exact scope of any waiver or agreement stemming from the confidential settlement is unknown, but without a merits-based decision, a new challenger would theoretically still be able to raise similar prior art grounds if they chose to file a new IPR.

There are no apparent pattern signals such as multiple IPRs from the same petitioner or aggressive appeals by the patent owner. The involvement of Perplexity AI, Inc. as a petitioner suggests a potential interest from an entity in the AI/ML space, but the settlement prevents drawing conclusions about the patent's susceptibility to challenge.

Recommended next steps

Since IPR2025-01063 was terminated by settlement, there is no Final Written Decision to link to or quote. For specific details regarding the IPR filing, such as the petition grounds and claims challenged, one would typically access the full case file through the USPTO Patent Trial and Appeal Board End-to-End (PTAB E2E) system.

For a defendant currently facing assertion of US11650968, the fact that the patent claims have not been adjudicated for validity by the PTAB means that any infringement theories are currently based on claims that have not been vetted through an inter partes proceeding. A potential defendant could consider initiating a new IPR if they have compelling prior art and a strong invalidity theory, as the previous settlement does not create a strong estoppel barrier against others. However, the reasons for the prior settlement should be investigated if possible, as they may offer insight into the patent owner's strategy or the perceived strength of the prior art.

Generated 5/16/2026, 6:47:54 PM

Ownership chain (1)

Asserters network →

Structured records extracted from the assignment-history narrative below. Each entity links to its full ownership-network profile.

  1. 2020-10-09 · reel 056461/0815 · Assignment of Assignors Interest

    LAHAV, NIMROD; MENDELS, GIDEON; NAIR, DHRUVComet ML, Inc.

    Correspondent: Jeffrey M. Hayes · Hayes Soloway

    internal reorg

Assignment history

Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.

✓ Generated

Inventors

  • Dhruv Nair (Comet Ml Inc)
  • Gideon Mendels (Comet Ml Inc)
  • Nimrod Lahav (Comet Ml Inc)

No unusual patterns noted; all inventors are associated with the original assignee.

Original assignee

Comet Ml Inc. is the entity named on the issued patent. Comet ML, Inc. provides a machine learning platform for tracking, comparing, and optimizing experiments, models, and datasets. They ship a product embodying the claims, specifically a platform that helps manage the machine learning lifecycle, which includes features for early stopping of neural network training. Comet Ml Inc. is currently operating.

Assignment timeline

  • 2020-10-09 (executed) / recorded 2020-10-09 — Reel 056461/0815
    • Conveyance: Assignment of Assignors Interest
    • Assignor: LAHAV, NIMROD; MENDELS, GIDEON; NAIR, DHRUV
    • Assignee: Comet ML, Inc.
    • Correspondent: Jeffrey M. Hayes, Esq., Hayes Soloway PC, 1048 Irvine Avenue #736, Newport Beach, CA 92660.
    • Context: Internal reorg – formal assignment of inventor rights to the company.

Timeline diagram

timeline
    title Ownership of US 11650968
    2019 : Application filed by Comet Ml Inc
    2020 : Inventors assigned to Comet ML Inc
    2023 : Patent granted

NPE / troll-pattern signals

  1. Shell-entity transfernot present. The initial assignment is from the inventors to Comet ML, Inc., which is an operating company.

  2. Known asserter in the chainnot present. Comet ML, Inc. is not a known NPE.

  3. Repeat correspondent across the chainnot present. Only one assignment is recorded, so no recurrence can be observed.

  4. Cascading transfersnot present. Only one assignment is recorded.

  5. Pre-litigation transferunclear. While the patent family has litigation (IPR2025-01063 was filed), there is only one recorded assignment which occurred before the patent issued and long before the reported litigation. Therefore, there's no evidence of a pre-litigation transfer in the assignment records themselves.

  6. Bankruptcy fire-salenot present. There is no indication that Comet ML, Inc. has filed for bankruptcy.

  7. Privateeringnot present. No evidence of transfer to an NPE for assertion on behalf of an operating company.

  8. Defensive aggregator (anti-NPE)not present. The patent is owned by an operating company, not a defensive aggregator.

Verdict

Insufficient data. Only one assignment from the inventors to the original assignee is recorded in the USPTO Assignment Center (reel 056461/0815). While Google Patents mentions an IPR case (IPR2025-01063), the USPTO assignment records do not show any subsequent transfers to an NPE or any other entity that would indicate a change in ownership for assertion purposes.

USPTO Assignment Center search page: https://assignmentcenter.uspto.gov/ (Search for patent number 11650968).

Generated 5/16/2026, 6:47:42 PM

Prior art

Earlier patents, publications, and products that may anticipate or render the claims unpatentable.

✓ Generated

The USPTO website provides tools to search for patents and their associated information. I have successfully accessed the patent text for US11650968B2 and will now proceed to identify the most relevant prior art.

To identify the most relevant prior art, I will examine the citations listed within US11650968B2 itself. These are typically the references that the patent examiner and the applicant considered most pertinent during prosecution.

Here's an analysis of the prior art cited in US11650968B2:

Prior Art Keywords from US11650968B2:

  • loss
  • training
  • model
  • threshold
  • data

I will now extract the specific patent and non-patent literature citations directly from the US11650968B2 patent document and analyze them.


Cited Patent References (US11650968B2):

The patent document lists various "Other References" and "U.S. Patent Documents." I will focus on the U.S. Patent Documents as they are generally considered direct prior art under 35 U.S.C. § 102.

US Patent Documents:

  1. US20170364969A1 (Published: 2017-12-21)

    • Inventors: Joshua B. Tenenbaum, Kevin L. Smith, Julian J. S. Schwarting
    • Assignee: Google LLC
    • Brief Description: This patent application describes methods and systems for training neural networks using adaptive optimization algorithms, which can include dynamically adjusting learning rates and other hyperparameters during training based on various metrics, including loss.
    • Potential Anticipation: This reference could potentially anticipate aspects of claims related to dynamically adjusting training parameters or stopping training based on optimization metrics (like loss), particularly if the "probability of improvement" or "wait value" described in Claim 1 could be construed as a form of adaptive optimization or dynamic stopping criterion based on observed training progress. For instance, the general concept of using training data to influence subsequent training decisions and resource allocation might be considered.
  2. US20180276632A1 (Published: 2018-09-27)

    • Inventors: Alexey Dosovitskiy, Philipp Fischer, Tobias Scheffer, Thomas Brox
    • Assignee: Google LLC
    • Brief Description: This patent application is related to training generative adversarial networks (GANs) and includes techniques for monitoring and controlling the training process, potentially involving early stopping or adjustments based on performance metrics to prevent issues like mode collapse or to optimize training efficiency.
    • Potential Anticipation: Similar to US20170364969A1, this reference might anticipate elements of Claim 1 regarding stopping training based on observed performance metrics and efficiency. The idea of monitoring training progress and making decisions to cease training prematurely to avoid undesirable outcomes or resource waste could be relevant to the inventive step of US11650968B2, especially concerning how "improvement" is measured or predicted.
  3. US20190244199A1 (Published: 2019-08-08)

    • Inventors: Justin Johnson, Alexandre F. Lacoste, Kyle H. S. Chard, Stephen L. W. Chan, Benjamin J. L. R. C. D. Van Der Ploeg, Matthew J. A. V. S. Hoffman, David J. L. M. Duvenaud
    • Assignee: Google LLC
    • Brief Description: This application discusses techniques for efficient hyperparameter optimization and neural architecture search, which inherently involve training multiple neural networks with different configurations and evaluating their performance. Early termination of unpromising configurations is a common practice in this field.
    • Potential Anticipation: This reference is highly relevant to Claim 1 and Claim 10 (and Claim 16 by extension) in the "global mode" context of US11650968B2, where hyperparameters are being chosen and training of NNs with different configurations is stopped. The concept of identifying when a particular NN with a given set of hyperparameters is unlikely to yield better results than others already evaluated, and subsequently stopping its training, directly addresses the core of the global early stopping strategy. The "probability that the likely best loss of the NN is lower than the best loss of the other NNs" in US11650968B2's claim 1 could be considered an advanced mechanism for making the decision to stop training an unpromising hyperparameter configuration, which is a goal also addressed by this prior art.

Note on Anticipation: Determining full anticipation under 35 U.S.C. § 102 requires a detailed claim-by-claim analysis, comparing each limitation of a claim to the disclosure of the prior art reference. The descriptions above provide a high-level assessment of potential relevance. A definitive conclusion would necessitate mapping each element of, for example, Claim 1, to explicit or inherent disclosures in these prior art documents. Specifically, the "model trained using training data from other NNs" to determine a probability of improvement or a likely best loss might be a distinguishing feature of US11650968B2 over more general early stopping or hyperparameter optimization techniques.

Generated 5/16/2026, 6:47:47 PM

Obviousness

Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.

✓ Generated

The following analysis addresses the obviousness of US Patent 11650968 under 35 U.S.C. § 103, identifying combinations of prior art references that would render the claims obvious and explaining the motivation for a person having ordinary skill in the art (POSITA) to combine them. The analysis primarily focuses on Independent Claim 1, as Claims 10 and 16 are directed to a system and a computer program product for carrying out the method of Claim 1, respectively.

The patent's priority date is May 24, 2019. The provided patent text describes the general state of prior art for neural network training: "Prior training methods typically stop by periodically testing a NN using a holdout data set not included in training data. Training may be stopped for example when improvement stagnates." [cite: US11650968 Description] The patent identifies a problem: "It is hard or impossible using prior art methods to predict at any given point in training how much improvement can be achieved by training using further epochs." [cite: US11650968 Description]

The references from the Google Patents listing for US11650968B2 (found under the "References" section) are used for this analysis.

Combination of Prior Art References

A combination of WO2019053350A1 ("EFFICIENT MACHINE LEARNING MODEL TRAINING WITH EARLY STOPPING" by Li et al., published March 21, 2019) and general knowledge within the field of machine learning, possibly supplemented by Prechelt, Lutz, "Early stopping—But when?" (2012), would render Claim 1 of US11650968 obvious.

WO2019053350A1 (Li et al.)
This international patent application, published prior to the priority date of US11650968, discloses methods for optimizing machine learning model training. Its abstract states: "Methods may include collecting metrics associated with a machine learning model training, determining an expected performance curve for the machine learning model training based on at least one trained model and the collected metrics, and stopping the machine learning model training based on the expected performance curve." [cite: WO2019053350A1 Abstract] The description further clarifies that "The trained model may be a surrogate model trained to predict future performance of a machine learning model based on its past performance." [cite: WO2019053350A1 Description]

Prechelt, Lutz, "Early stopping—But when?" (2012)
This seminal work discusses various strategies for early stopping based on monitoring validation error to prevent overfitting during neural network training. It highlights the reactive nature of conventional early stopping methods, which cease training when observed improvement stagnates. [cite: Prechelt (2012)]

General Knowledge of a Person Having Ordinary Skill in the Art (POSITA)
A POSITA in machine learning as of 2019 would possess knowledge of:

  • Standard neural network training processes, including iterative epochs and loss computation.
  • The computational expense of training large neural networks and the desirability of early stopping.
  • Techniques for training predictive models (e.g., surrogate models) on diverse datasets to achieve robustness and generalization.
  • Basic statistical methods for deriving probabilities and expected values from predictive model outputs.
  • Using thresholds for decision-making in automated processes.

Obviousness Analysis of Claim 1

Claim 1: A computer-implemented method for predictively stopping training of a neural network (NN), the method comprising:

1a. training, by one or more processors, the NN over a series of training periods, wherein a loss for the NN is computed during each training period;

  • Disclosure in Prior Art: This element is widely known in the art and explicitly taught by both Prechelt (2012) and WO2019053350A1. Training neural networks iteratively and computing a loss function are fundamental aspects of machine learning. WO2019053350A1 describes "collecting metrics associated with a machine learning model training," which includes computing loss during training periods. [cite: WO2019053350A1 Abstract]

1b. determining, by the one or more processors and using one or more models that have been trained using training losses of a plurality of NNs other than the NN, a probability of improvement in the loss of the NN, wherein the one or more models have input thereto a set of model parameters for the NN and data describing training loss of the NN;

  • Disclosure in Prior Art: WO2019053350A1 discloses "determining an expected performance curve for the machine learning model training based on at least one trained model and the collected metrics." [cite: WO2019053350A1 Abstract] It further states that this "trained model may be a surrogate model trained to predict future performance of a machine learning model based on its past performance." [cite: WO2019053350A1 Description]
    • "using one or more models that have been trained using training losses of a plurality of NNs other than the NN": While WO2019053350A1 doesn't explicitly use the phrase "plurality of NNs other than the NN," a POSITA would understand that a "surrogate model trained to predict future performance of a machine learning model" would inherently be trained on a diverse dataset of training histories from many different machine learning models or configurations. Training such a predictive model on data from only a single NN would limit its generalizability, which is contrary to the purpose of a surrogate model designed for "efficient machine learning model training." Therefore, it would be an obvious design choice for a POSITA to train the "at least one trained model" using data from a plurality of NNs to make it broadly applicable and robust.
    • "a probability of improvement in the loss of the NN": Once an "expected performance curve" or "predicted future performance" is determined (as taught by WO2019053350A1), calculating a "probability of improvement" (e.g., the likelihood that the loss will fall below a certain target value or continue to decrease by a significant amount) is a routine statistical derivation. US11650968 itself describes using a cumulative distribution function (CDF) with a mean and variance from a model to determine such a probability. [cite: US11650968 Description] This is a standard application of predictive analytics.
    • "input thereto a set of model parameters for the NN and data describing training loss of the NN": WO2019053350A1 describes using "collected metrics" to determine the expected performance curve, and specifically mentions the surrogate model predicting future performance "based on its past performance." [cite: WO2019053350A1 Abstract, WO2019053350A1 Description] This directly corresponds to inputting parameters and loss history.

1c. stopping, by the one or more processors, the training of the NN if the determined probability is less than a probability threshold, or if a wait value is greater than a wait threshold.

  • Disclosure in Prior Art: WO2019053350A1 teaches "stopping the machine learning model training based on the expected performance curve." [cite: WO2019053350A1 Abstract] The use of "thresholds" for decision-making based on a calculated metric (such as a probability) is a fundamental engineering principle. Given the derivation of a "probability of improvement," it would be obvious to a POSITA to define a "probability threshold" to trigger early stopping. The "wait value" (or patience) mechanism, which stops training after a period of non-improvement, is a well-established heuristic in early stopping, as taught by Prechelt (2012) and generally known in the art. [cite: Prechelt (2012)]

Motivation to Combine

A POSITA would be motivated to combine these references and general knowledge for the following reasons:

  1. Addressing Computational Inefficiency: The core problem addressed by US11650968—the wasteful expenditure of computational resources on neural network training when improvement is no longer likely—was a well-known challenge in the art. Prior art like Prechelt (2012) highlighted the importance of early stopping but primarily offered reactive solutions based on observed stagnation. A strong motivation existed to develop more proactive and efficient stopping mechanisms. [cite: US11650968 Description]
  2. Improving Stopping Accuracy and Efficiency: WO2019053350A1 offered a significant step forward by introducing a "trained model" to predict an "expected performance curve." A POSITA would be motivated to integrate this predictive capability into existing early stopping strategies to achieve more precise and timely cessation of training.
  3. Generalization for Robustness: To make a predictive early stopping model broadly useful across different neural network architectures and hyperparameters, a POSITA would naturally train the "surrogate model" (from WO2019053350A1) on a diverse dataset of training runs from "a plurality of NNs." This is a standard approach in machine learning to ensure robustness and generalizability, thereby making the predictive model more valuable and widely applicable.
  4. Deriving Actionable Metrics: Once a predicted performance curve is available, it would be a straightforward and obvious step for a POSITA to derive actionable metrics such as the "probability of improvement." Such probabilities provide a clear, quantifiable basis for making stopping decisions, enhancing the sophistication of the early stopping mechanism.
  5. Integrating Established Heuristics: Combining a predictive probability threshold with a traditional "wait value" (patience) mechanism from existing early stopping methods (like those in Prechelt (2012)) would be an obvious way to create a robust and comprehensive stopping policy, ensuring that training is halted both when improvement is predicted to be unlikely and when it simply fails to materialize over time.

In conclusion, the combination of the teachings of WO2019053350A1, the well-known early stopping techniques described in works like Prechelt (2012), and the general understanding of machine learning principles by a POSITA would have made the method of Claim 1 of US11650968 obvious at the time of invention. The advancements claimed in US11650968 represent predictable developments and applications of known techniques to address a recognized problem in the field.

Generated 5/16/2026, 6:48:31 PM

Extensions

Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.

✓ Generated

To provide a comprehensive analysis of US Patent 11650968, I will detail any patent term adjustments (PTA), patent term extensions (PTE), continuation applications, divisional applications, related family members, and the projected expiration date, based on the current date of April 26, 2026.

Patent Term Adjustments (PTA)

Patent Term Adjustment (PTA) is granted to a patent to compensate for certain delays caused by the USPTO during the prosecution of a utility or plant patent application. This adds time to the standard 20-year patent term from the filing date. Key USPTO delays that trigger PTA include:

  • Failure to issue a first office action or notice of allowance within 14 months of filing.
  • Failure to act on an application within four months of an applicant's response to an office action.
  • Failure to issue a patent within four months of payment of the issue fee.

Applicant delays can reduce or eliminate PTA.

The USPTO automatically determines the PTA period and transmits a notice of this determination no later than the patent's issue date. Applicants have one opportunity to request reconsideration of the PTA calculation within two months of the patent's issuance, a deadline that can be extended by up to five months with appropriate fees.

To determine the specific PTA for US11650968, direct access to the patent's prosecution history in USPTO Patent Center or Public Search is required, which is beyond the scope of this tool. However, the Google Patents entry for US11650968B2 indicates an "Adjusted expiration" date of 2041-10-22, which suggests that some form of patent term adjustment has already been calculated and applied, extending it beyond the typical 20 years from its filing date of 2019-12-09.

Patent Term Extensions (PTE)

Patent Term Extensions (PTE) are distinct from PTA and are available for patents claiming certain human drug products, medical device products, animal drug products, veterinary biological products, and food or color additive products. PTE aims to restore patent term lost due to premarket government approval processes by regulatory agencies like the FDA. The extension is limited to a maximum of five years, and the total patent life with a PTE cannot exceed 14 years from the date of FDA approval.

Since US11650968 relates to "Systems and methods for predictive early stopping in neural network training" and not to a regulated product requiring premarket approval, it is highly unlikely to be eligible for a Patent Term Extension under 35 U.S.C. 156.

Continuation and Divisional Applications

  • Continuation Applications: A continuation application uses the same disclosure as an earlier non-provisional application, filed while the earlier application is still pending. It allows for further examination of claims not allowed in the parent application.
  • Divisional Applications: A divisional application is filed when an original application claims two or more independent and distinct inventions, and the USPTO requires restriction to one invention. The other invention(s) can be pursued in a divisional application, retaining the benefit of the original filing date.
  • Continuation-in-Part (CIP) Applications: A CIP application includes subject matter from an earlier-filed application and also introduces new matter not disclosed in the earlier application.

The Google Patents entry for US11650968B2 lists "US20200372342A1" as another version. This typically indicates a published patent application. To definitively determine if US20200372342A1 is a continuation, divisional, or related application, one would need to review the application's file wrapper in USPTO Patent Center or Public Search. Based on the information provided, it is a related family member, likely the published application that led to the granted patent.

Related Family Members

The patent family for US11650968B2, as provided by Google Patents, includes:

  • US16/707,265 (Application number)
  • US201916707265A (another format of the application number)
  • US20200372342A1 (Publication number of the application)

The relationship between these is that US16/707,265 is the application number, US20200372342A1 is the publication of that application, and US11650968B2 is the granted patent resulting from that application.

Projected Expiration Date

The "Adjusted expiration" date provided on Google Patents for US11650968B2 is 2041-10-22. This date already incorporates any Patent Term Adjustment (PTA) applied to the patent. Given that the patent is not related to a product subject to regulatory approval, Patent Term Extension (PTE) is not applicable. Therefore, the projected expiration date is 2041-10-22.

Generated 5/16/2026, 6:47:57 PM

Derivative works

Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.

✓ Generated

Defensive Disclosure for US Patent 11650968: Systems and Methods for Predictive Early Stopping in Neural Network Training

This document details derivative variations and integration scenarios for US Patent 11650968, aiming to broaden the scope of disclosed prior art and render incremental improvements by competitors "obvious" or "non-novel" to a person having ordinary skill in the art (PHOSITA). These disclosures are based on the core inventive concepts of predictively stopping neural network (NN) training using a model trained on other NNs' loss data to determine a probability of improvement.


Derivative 1.1: Material & Component Substitution - Neuromorphic Hardware for Training and Early Stopping Model Execution

Enabling Description: The computer-implemented method for predictively stopping neural network training is instantiated on a neuromorphic computing system. Specifically, the target neural network (NN) is a Spiking Neural Network (SNN) trained directly on a neuromorphic processing unit (NPU), such as an IBM TrueNorth or Intel Loihi chip. Training involves adjusting synaptic weights based on spike-timing-dependent plasticity (STDP) or other neuromorphic learning rules. The "loss" for each training period (epoch) is characterized by aggregated event-driven metrics, such as a deviation in output spike patterns from target patterns, total energy consumption per inference, or a measure of network entropy. This event-stream loss history, along with the SNN's topological parameters, is transmitted to a dedicated, co-located neuromorphic early stopping predictive model. This predictive model, also implemented as an SNN on the NPU and pre-trained on a corpus of event-stream loss histories from various other SNN training runs, determines the probability of future improvement. The probability calculation leverages sparse, event-based computations. If the determined probability falls below a pre-configured spike-rate threshold, or if an event-count-based wait value exceeds its limit, a halt signal is issued across the NPU's inter-chip communication fabric, pausing or terminating the SNN training to conserve neuromorphic computational cycles and power.

flowchart TD
    A[SNN Training on Neuromorphic Substrate] --> B{Compute Spike-Rate Loss / Energy Metrics};
    B --> C[Generate Event-Stream Loss History];
    C --> D[Neuromorphic Early Stopping Model (Predictive Model)];
    D -- Inputs: SNN Parameters, Event-Stream Loss History --> E{Determine Probability of Improvement};
    E -- Prob < Threshold OR Wait > Threshold --> F{Halt Signal via Inter-Chip Fabric};
    E -- Else --> A;

Derivative 1.2: Material & Component Substitution - FPGA-accelerated Early Stopping Model with Custom Loss Accumulators

Enabling Description: In this derivative, the training of the target neural network is performed on conventional Graphics Processing Unit (GPU) clusters. However, the predictive early stopping model is accelerated and implemented on a Field-Programmable Gate Array (FPGA). Loss data from the GPU-based NN training, typically a floating-point scalar per epoch, is streamed over a high-bandwidth interconnect (e.g., PCIe, NVLink) to the FPGA. On the FPGA, custom hardware logic gates are configured to perform real-time feature extraction from the incoming loss history, including calculating first and second-order differences, running-mean, and standard deviation in fixed-point arithmetic. The core of the predictive model, such as a Gradient Boosting Machine (GBM) decision tree ensemble, is instantiated with optimized hardware blocks for ultra-low latency inference. The FPGA's digital logic directly performs the probabilistic determination and threshold comparisons. Upon satisfaction of the stopping criteria (e.g., probability below a configurable threshold or wait value exceeding a maximum count), the FPGA generates a dedicated control signal that is sent directly to the GPU cluster's scheduler, issuing a command to halt or suspend the associated NN training job, thereby offloading critical decision-making from the host CPU and achieving microsecond-level response times.

graph TD
    A[GPU Cluster NN Training] --> B{Stream Loss Data};
    B --> C[FPGA Early Stopping Accelerator];
    C -- Custom Logic: GBM Inference, Feature Extraction, Prob Calc --> D{Decision Logic};
    D -- Stop Training --> E[Control Signal to GPU Cluster];
    D -- Continue Training --> B;

Derivative 1.3: Material & Component Substitution - In-Memory Database for Loss History and Federated Learning Model for Prediction

Enabling Description: This derivative employs a distributed in-memory database system (e.g., Apache Ignite, Redis Enterprise) for managing the extensive loss histories of all neural networks, both the target NN and the plurality of other NNs used to train the predictive model. This architecture minimizes I/O latency, facilitating rapid access to historical data. The predictive early stopping model itself is realized through a federated learning paradigm. Multiple client nodes, each responsible for training a subset of NNs or specific hyperparameter configurations, locally maintain and update instances of the early stopping predictive model. These local models are trained on loss histories pertaining to their respective training tasks, leveraging the low-latency in-memory data store. Periodically, these local models transmit anonymized model updates (e.g., gradient aggregates, differential privacy-preserved model parameters) to a central orchestrator. The orchestrator synthesizes these updates into a global early stopping model, which is then broadcast back to all client nodes. When a client trains a target NN, its local federated early stopping model, informed by the latest global model, queries the in-memory database for the target NN's loss history and parameters, computes the probability of improvement, and makes a local decision to stop or continue training based on predefined thresholds.

sequenceDiagram
    participant C1 as Client 1 (NN Training)
    participant C2 as Client 2 (NN Training)
    participant DB as In-Memory Database
    participant O as Orchestrator (Federated Learning)
    C1->>DB: Store Loss History (NN_A)
    C2->>DB: Store Loss History (NN_B)
    O->>DB: Aggregate Anonymized Loss Patterns
    O->>O: Refine Global Early Stopping Model
    O->>C1: Distribute Global Model Update
    C1->>C1: Update Local Federated ES Model
    C1->>DB: Query Loss History (NN_A)
    C1->>C1: Local ES Model Predicts Probability
    C1->>C1: Stop NN_A Training if criteria met

Derivative 1.4: Operational Parameter Expansion - Predictive Early Stopping for Exascale Foundation Model Training

Enabling Description: The computer-implemented method is adapted for the training of massively parameterized foundation models on exascale high-performance computing (HPC) systems. The "target NN" comprises billions to trillions of parameters, trained across hundreds of thousands of interconnected processing units. The "loss" function includes not only primary objective metrics (e.g., perplexity for language models, reconstruction error for generative models) but also secondary, emergent capability indicators derived from periodic evaluations on representative downstream tasks. The "training periods" are measured in highly granular, distributed computational steps rather than traditional epochs. The predictive early stopping model is itself a sophisticated, distributed meta-learning agent (e.g., a large-scale recurrent neural network or transformer) trained on an exabyte-scale dataset of historical training curves, intermediate checkpoint performance, and resource utilization profiles from prior exascale model development efforts. This meta-model accurately determines the probability of achieving further meaningful improvement in the target foundation model's performance within the current training trajectory. Dynamic thresholds are employed, adapting based on observed scaling laws and the current optimization phase (e.g., higher patience during critical emergent learning phases). Stopping signals are disseminated through a low-latency, resilient HPC messaging interface to prevent wasted exaflops and energy consumption.

graph TD
    A[Exascale Compute System] --> B[Distributed Foundation Model Training];
    B --> C{Collect Loss & Performance Metrics (Exabytes)};
    C --> D[Massive Transformer-based ES Model];
    D -- Inputs: Multi-modal training metrics, historic exascale runs --> E{Compute Prob. of Improvement (High Rigor)};
    E -- Prob < Dynamic Threshold OR Wait > Threshold --> F{Distributed Halt Signal};
    F --> G[Optimize Resource Allocation];
    E -- Else --> B;

Derivative 1.5: Operational Parameter Expansion - Micro-Scale Edge AI Training with Resource-Constrained Early Stopping

Enabling Description: This derivative applies predictive early stopping to neural networks undergoing continuous, incremental training on resource-constrained edge devices, such as microcontrollers or embedded System-on-Chips (SoCs). The "target NN" is typically a highly quantized, compact model for tasks like keyword spotting or anomaly detection, constantly adapting to local sensor data. Due to severe memory and computational limitations (e.g., KB of RAM, MHz clock speeds), the "loss history" is maintained as a compact, fixed-size circular buffer storing the last N validation loss values. The "model trained using training losses of a plurality of NNs" is a pre-trained, deeply compressed model (e.g., an extremely shallow decision tree ensemble or a quantized linear regression model), which is executed directly on the edge device's low-power microcontroller unit. Feature extraction from the loss history is minimalistic (e.g., simple moving average, last derivative). The "probability of improvement" calculation is approximated using integer or fixed-point arithmetic. The stopping criteria are critically sensitive to power consumption: training is halted not only when performance gains plateau, but also if the marginal improvement per Joule consumed falls below an energy-efficiency threshold. This ensures maximal battery life and sustained on-device operation.

stateDiagram-v2
    state "Edge Device Idle" as Idle
    state "Collect Sensor Data" as Collect
    state "Tiny NN Training" as TrainNN
    state "Compute Quantized Loss" as ComputeLoss
    state "Update Rolling Loss History" as UpdateHistory
    state "Run Lightweight ES Model" as RunES
    state "Determine Stop Condition" as DecideStop

    Idle --> Collect: Sensor Event
    Collect --> TrainNN: New Data Batch
    TrainNN --> ComputeLoss: Epoch Complete
    ComputeLoss --> UpdateHistory: New Loss Value
    UpdateHistory --> RunES: Loss History Ready
    RunES --> DecideStop: Predictive Output
    DecideStop -- Prob < Threshold OR Wait > Threshold --> Idle: Stop Training (Save Energy)
    DecideStop -- Else --> TrainNN: Continue Training

Derivative 1.6: Cross-Domain Application - Predictive Early Stopping for Biopharmaceutical Drug Discovery (Protein Folding Models)

Enabling Description: The predictive early stopping method is deployed within the computational pipeline for biopharmaceutical drug discovery, specifically for training deep learning models that predict protein folding or simulate molecular dynamics. The "target NN" is a complex graph neural network, a transformer-based architecture (e.g., similar to AlphaFold), or a large-scale generative model, configured to predict the 3D structure of novel proteins from amino acid sequences, or to estimate binding affinities of potential drug candidates. The "loss" function integrates metrics such as Root Mean Square Deviation (RMSD) from experimentally determined structures (if available), predicted binding free energy (ΔG), and metrics quantifying conformational stability or diversity. The "model trained using training losses of a plurality of NNs" is a specialized ensemble predictive model trained on a vast historical dataset comprising thousands of previous protein folding and molecular simulation model training runs. This historical data includes loss curves, convergence patterns, and final RMSD/ΔG scores across a diversity of protein targets, ligand chemistries, and model hyperparameters. This predictive model determines the probability that further training epochs for a current target NN will yield significant improvements in structural accuracy or binding affinity. Early stopping prevents over-allocation of supercomputing resources to models whose potential for further biological relevance is statistically low.

flowchart LR
    A[Raw Protein Sequence Data] --> B{Pre-processing & Graph Generation};
    B --> C[Target Graph NN (Protein Folding)];
    C -- Training Epochs --> D{Compute RMSD & Binding Loss};
    D --> E[Record Loss History & Model State];
    E --> F[Predictive Early Stopping Model (Trained on Protein Folding History)];
    F -- Inputs: NN State, Loss History, Protein Features --> G{Determine Prob. of Improvement (RMSD/Binding)};
    G -- Prob < Threshold OR Wait > Threshold --> H[Halt Supercomputer Training Job];
    G -- Else --> C;

Derivative 1.7: Cross-Domain Application - Predictive Early Stopping for Autonomous Navigation System Training (Reinforcement Learning)

Enabling Description: This derivative applies the early stopping methodology to the training of reinforcement learning (RL) agents for autonomous navigation within high-fidelity simulation environments (e.g., for self-driving vehicles, robotic manipulators, or aerial drones). The "target NN" is the policy network (e.g., a Deep Q-Network, Actor-Critic, or PPO agent) of the RL system. The "loss" is derived from the negative of cumulative episode rewards, policy entropy, or value function errors. The "training periods" correspond to blocks of simulation episodes. The "model trained using training losses of a plurality of NNs" is a meta-RL model or a statistical surrogate model. It is trained on a comprehensive historical dataset of RL agent training logs, including episode rewards, convergence rates, exploration metrics, and final task performance (e.g., success rate, collision avoidance) across various RL algorithms, environmental complexities, and hyperparameter settings. This predictive model determines the probability that the current RL agent will achieve a higher cumulative reward, reduced collision frequency, or improved task completion rate beyond its current trajectory, or surpass the performance of other candidate RL agents. Early stopping mitigates simulation resource expenditure and prevents over-optimization to simulator artifacts, facilitating faster deployment of robust policies.

sequenceDiagram
    participant S as Simulation Environment
    participant RL as RL Agent (Target NN)
    participant ES as Early Stopping Model
    S->>RL: State Observation
    RL->>S: Action
    S->>RL: Reward, Next State
    loop Training Episode
        RL->>RL: Update Policy Network (compute "loss")
        RL->>RL: Record Episode Rewards/Metrics
    end
    RL->>ES: Send Training Metrics (Loss History, Rewards)
    ES->>ES: Determine Prob. of Improvement (Cumulative Reward)
    ES-->>RL: Stop Signal if (Prob < Threshold OR Wait > Threshold)
    RL->>RL: Halt Training or Continue

Derivative 1.8: Cross-Domain Application - Predictive Early Stopping for Predictive Maintenance Models in Industrial IoT

Enabling Description: The early stopping method is implemented for the training of predictive maintenance machine learning models in an Industrial Internet of Things (IIoT) ecosystem. The "target NN" is a recurrent neural network (RNN), a temporal convolutional network (TCN), or a transformer model, designed to process high-frequency time-series data from industrial sensors (e.g., vibration, temperature, current, acoustic emissions) to forecast impending equipment failures. The "loss" function is tailored for prognostics, incorporating metrics such as mean absolute error (MAE) in Remaining Useful Life (RUL) prediction, F1-score for fault classification, or economic cost of false alarms/missed detections. The "model trained using training losses of a plurality of NNs" is a domain-specific predictive model, leveraging a repository of historical training curves, hyperparameter configurations, and deployment performance from various predictive maintenance projects across different industrial assets (e.g., pumps, motors, turbines) and manufacturing plants. This predictive model quantifies the probability that the current RNN's RUL prediction accuracy or fault detection capability will improve beyond its current state, or outperform other models. Early stopping reduces compute load on distributed IIoT training infrastructure (e.g., fog computing nodes) and ensures that effective models are deployed to prevent costly equipment downtime without prolonged, inefficient training cycles.

flowchart TD
    A[Industrial Machinery Sensors] --> B[Time Series Data Stream (Vibration, Temp)];
    B --> C[Data Preprocessing (Anomaly Detection, Feature Eng)];
    C --> D[Target RNN/Transformer (Failure Prediction)];
    D -- Training Epochs --> E{Compute Prediction Accuracy Loss (FPs/FNs)};
    E --> F[Record Loss History];
    F --> G[Predictive Early Stopping Model (Trained on IIoT ML History)];
    G -- Inputs: RNN Params, Loss History, Machine Type --> H{Determine Prob. of Improvement (Accuracy/Cost)};
    H -- Prob < Threshold OR Wait > Threshold --> I[Halt Training & Deploy];
    H -- Else --> D;

Derivative 1.9: Integration with Emerging Tech - AI-driven Adaptive Thresholding for Predictive Early Stopping

Enabling Description: The predictive early stopping method is enhanced by an autonomous AI-driven adaptive thresholding agent. Instead of static, pre-configured probability and wait thresholds, these critical parameters (e.g., probability_threshold, wait_threshold) are dynamically adjusted in real-time by a meta-learning system. This meta-agent, which could be another reinforcement learning agent, a Bayesian optimization system, or an evolutionary algorithm, continuously monitors the efficacy of the early stopping mechanism itself. It collects feedback on various performance indicators such as total training time saved, final model generalization error relative to theoretical optimum, computational cost overhead of the early stopping process, and instances of premature stopping or late stopping. Based on this meta-data, the AI agent iteratively refines the thresholds, predicting optimal values for the current target NN's training context (e.g., specific dataset, architecture, available computational budget, and business objectives prioritizing speed vs. absolute accuracy). This creates a self-optimizing early stopping system that intelligently adapts its stopping policy to achieve superior overall resource efficiency and model development outcomes across diverse machine learning workflows.

stateDiagram-v2
    state "Initialize Fixed Thresholds" as Init
    state "NN Training" as TrainNN
    state "Early Stopping Model Prediction" as ES_Predict
    state "Apply Stopping Logic" as ApplyStop
    state "Collect ES System Performance Metrics" as CollectMetrics
    state "AI-driven Adaptive Threshold Agent" as AdaptiveAgent
    state "Update Adaptive Thresholds" as UpdateThresholds

    Init --> TrainNN
    TrainNN --> ES_Predict: Loss History, NN Params
    ES_Predict --> ApplyStop: Prob, Wait
    ApplyStop --> TrainNN: Continue (if not stopped)
    ApplyStop --> CollectMetrics: Stop/Continue Decision & Outcomes
    CollectMetrics --> AdaptiveAgent: ES Performance Data
    AdaptiveAgent --> UpdateThresholds: New Optimal Thresholds
    UpdateThresholds --> TrainNN: Apply New Thresholds

Derivative 1.10: Integration with Emerging Tech - Blockchain for Verifiable Loss Histories and Auditable Early Stopping

Enabling Description: This derivative integrates blockchain technology to establish an immutable and auditable record of neural network training progress and early stopping decisions. For each training period (epoch) of a target NN, the computed loss, associated hyperparameters, a cryptographic hash of the training data subset used, and the current state of the NN (e.g., model weights hash) are packaged into a transaction. This transaction is cryptographically signed and appended to a distributed ledger (blockchain). This creates a tamper-proof "loss history" record for the NN. The parameters and training data references of the "model trained using training losses of a plurality of NNs" (e.g., the LightGBM models) are also recorded on the blockchain, ensuring transparency of the predictive mechanism itself. When the early stopping model determines a probability of improvement and subsequently issues a stop/continue decision, this decision, along with the inputs (loss history hash, current hyperparameters), the calculated probability, and the thresholds applied, is also recorded as a signed transaction on the blockchain. This provides a verifiable, unalterable audit trail for every training run and stopping decision, which is critical for compliance in regulated industries (e.g., financial services, healthcare) and for establishing trusted provenance of AI models. Smart contracts can be deployed to automatically enforce stopping rules or trigger alerts based on on-chain conditions.

flowchart TD
    A[NN Training Epoch] --> B{Compute Loss, Collect Params};
    B --> C[Hash Loss & Params];
    C --> D[Create Blockchain Transaction];
    D -- Add to Ledger (Immutable) --> E[Distributed Ledger (Blockchain)];
    E --> F{Retrieve Verifiable Loss Histories & Model Params};
    F --> G[Predictive Early Stopping Model (Oracle Node)];
    G -- Inputs: Verifiable History, Current Params --> H{Determine Prob. of Improvement};
    H -- Decision (Stop/Continue) --> I[Record Decision on Blockchain (Signed)];
    I --> J[Auditable Training Record];

Derivative 1.11: Integration with Emerging Tech - Real-time IoT Performance Monitoring for Early Stopping with Telemetry

Enabling Description: The early stopping method is augmented by integrating real-time operational performance telemetry from deployed IoT devices that utilize the neural network post-training. For instance, a NN trained for computer vision tasks (e.g., object detection) on smart cameras, has its training progress informed by live feedback loops from a fleet of deployed cameras. This real-world telemetry includes metrics such as actual detection accuracy on live data streams, inference latency under varying network conditions, false positive/negative rates in different environments, and energy consumption during real-world inference. This live, contextual IoT data is streamed back to the training system. The "loss history" for the predictive early stopping model is dynamically extended to include these real-world performance indicators, alongside traditional validation loss. Consequently, the "probability of improvement" is refined to represent the likelihood of achieving better real-world operational performance (e.g., higher true positive rate, lower inference energy) rather than solely optimizing for static validation set metrics. This approach enables a more robust and practical early stopping mechanism, preventing models from over-training on potentially unrepresentative validation data and ensuring fitness for purpose in diverse operational environments.

graph TD
    subgraph Training System
        A[NN Training] --> B{Compute Validation Loss};
        B --> C[Predictive ES Model];
    end

    subgraph IoT Deployment
        D[Deployed NN on IoT Device] --> E{Real-time Inference};
        E --> F[Performance Telemetry (Accuracy, Latency)];
        F --> G[IoT Sensors (Contextual Data)];
    end

    F & G --> H[Telemetry Aggregation & Preprocessing];
    H --> C;
    C -- Inputs: Val Loss, IoT Telemetry, Context --> I{Determine Prob. Real-World Improvement};
    I -- Prob < Threshold OR Wait > Threshold --> A: Halt Training;
    I -- Else --> A: Continue Training;

Derivative 1.12: The "Inverse" or Failure Mode - Energy-Aware Early Stopping for Sustainable AI Training

Enabling Description: This derivative integrates explicit energy cost considerations into the early stopping decision process. The "computer-implemented method" tracks, alongside the traditional loss, the cumulative energy consumption (e.g., GPU Watt-hours, total data center power draw) associated with the target NN's training. The "model trained using training losses of a plurality of NNs" is extended to incorporate historical energy consumption profiles alongside loss histories. The "probability of improvement" is redefined as the likelihood of achieving a target loss within a predefined energy budget, or, more precisely, the probability of obtaining a significant marginal improvement in loss per unit of additional energy expended. A dynamic energy_efficiency_threshold is introduced: if the predicted energy cost to achieve further significant loss reduction exceeds this threshold, or if a dedicated energy_wait_value (incremented when energy efficiency drops below a minimum) surpasses its limit, training is stopped. This allows for a "low-power" or "sustainable AI" early stopping mode, where training is halted even if minor performance gains are still possible, prioritizing energy conservation and carbon footprint reduction over maximizing infinitesimal model performance.

flowchart TD
    A[NN Training] --> B{Compute Loss & Energy Consumption};
    B --> C[Loss History + Energy History];
    C --> D[Predictive ES Model (Energy-Aware)];
    D -- Inputs: NN Params, Loss/Energy History --> E{Determine Prob. of "Efficient" Improvement};
    E -- Prob < Energy_Threshold OR Wait > Energy_Wait_Threshold --> F[Halt Training (Energy Optimized)];
    E -- Else --> A;

Derivative 1.13: The "Inverse" or Failure Mode - Degradation-Tolerant Early Stopping for Continual Learning Systems

Enabling Description: The early stopping mechanism is specifically tailored for continual learning (CL) systems, where neural networks learn sequentially from non-stationary data streams without forgetting previously acquired knowledge. For the "target NN" in a CL setting, the "loss" function is augmented to include a "forgetting loss" (e.g., evaluation on a small, representative replay buffer or on a set of pseudo-rehearsal samples from old tasks), in addition to the current task loss. The "model trained using training losses of a plurality of NNs" is built from historical CL experiments, encompassing various CL strategies (e.g., regularization-based, replay-based) and their associated performance trajectories on both current and old tasks. This predictive model determines the "probability of net improvement," defined as the likelihood of achieving better performance on the current task while simultaneously avoiding catastrophic forgetting on previous tasks. If the predictive model forecasts that continued training on the new data stream will lead to an unacceptable probability of significant degradation on previously learned tasks (even if current task loss is improving), early stopping is triggered. This forces the CL system into a "degradation-avoidance" or "limited-functionality" mode, where it might stop learning, trigger a consolidation phase, or initiate a task-switching mechanism.

flowchart LR
    A[New Task Data Stream] --> B{Continual Learning NN Training};
    B -- Training Epochs --> C{Compute Current Task Loss & Forgetting Loss (Replay Buffer)};
    C --> D[Combined Loss History];
    D --> E[Predictive ES Model (Degradation-Aware)];
E -- Inputs: NN State, Combined Loss History --> F{Determine Prob. of NET Improvement (New + Old Tasks)};
    F -- Prob < Degradation_Threshold OR Wait > Forgetting_Wait --> G[Halt Training (Prevent Forgetting)];
    G -- Else --> B;

Derivative 1.14: The "Inverse" or Failure Mode - Safe-Failure Early Stopping for Critical Systems

Enabling Description: This derivative implements a safety-critical variant of the predictive early stopping method, designed for neural networks operating in high-assurance or life-critical applications (e.g., medical diagnostics, autonomous flight control, industrial safety systems). The "loss" function for the target NN is augmented with real-time safety metrics, such as a probabilistic quantification of system uncertainty, deviation from verified safe operating envelopes, or the predicted probability of violating critical safety constraints. The "model trained using training losses of a plurality of NNs" is explicitly trained on a curated dataset of safety-validated training runs, where each historical trajectory includes both performance loss and comprehensive safety assurance metrics. This specialized predictive model calculates the "probability of safe improvement," which is the likelihood that further training will enhance performance without compromising safety-critical thresholds. If the predictive model determines that the probability of maintaining safety during continued training falls below a stringent, pre-defined critical_safety_threshold, or if it forecasts an increasing trend toward an unsafe operating region, training is immediately halted in a "safe-failure mode." This mode prioritizes system integrity and human safety by, for example, reverting to the last-known safe model checkpoint, triggering an emergency alert, or initiating a controlled shutdown process, overriding traditional performance-only stopping criteria.

graph TD
    A[NN Training (Critical System)] --> B{Compute Standard Loss & Safety Metrics (e.g., Uncertainty, Safety Constraint Violation)};
    B --> C[Augmented Loss History (Includes Safety)];
    C --> D[Predictive ES Model (Safety-Critical)];
    D -- Inputs: NN State, Augmented Loss History, Safety Parameters --> E{Determine Prob. of SAFELY Improving};
    E -- Prob_Safety < Critical_Threshold --> F[Initiate Safe-Failure Protocol];
    F --> G[Halt Training & Revert to Safe State / Alert];
    E -- Else --> A;

Combination Prior Art Scenarios with Open-Source Standards

The core method of US11650968 can be combined with existing open-source standards, making further incremental improvements obvious to a person skilled in the art.

  1. US11650968 + MLflow Tracking Standard:

    • Description: The predictive early stopping method (Claim 1) is integrated into an experimental tracking workflow managed by the open-source MLflow platform. During the training of a target neural network, each epoch's loss and other relevant metrics are logged using mlflow.log_metric(). The loss history described in US11650968 is directly extracted from the MLflow tracking server's backend database (e.g., PostgreSQL, SQLite, or equivalent artifact store for mlruns). The "model trained using training losses of a plurality of NNs" ingests this MLflow-tracked data to build its predictive capabilities. When the early stopping criteria (probability < threshold OR wait value > threshold) are met, the early stopping module utilizes mlflow.end_run() or a custom status update to mark the experiment as stopped and log the reason, final predicted loss, and other decision parameters back into MLflow, creating a comprehensive and auditable record of the early stopping decision within a standard ML lifecycle management framework.
    • Obviousness: It would be obvious to a PHOSITA in MLOps to integrate an efficient early stopping mechanism with a widely adopted experiment tracking standard like MLflow. The core functionalities of MLflow—logging metrics, parameters, and managing run lifecycles—directly align with the data collection and control requirements for implementing and documenting the patented early stopping method.
  2. US11650968 + ONNX (Open Neural Network Exchange) Standard:

    • Description: The predictive early stopping model itself (e.g., a LightGBM ensemble or other tree-based model as mentioned in the patent) is converted and serialized into the Open Neural Network Exchange (ONNX) format. This ONNX-formatted early stopping model is then deployed via an ONNX Runtime. When the target NN is being trained, its loss history and hyperparameters are processed to extract the necessary features. These features are then provided as input to the ONNX Runtime, which executes the ONNX-formatted predictive early stopping model to determine the probability of improvement. This allows the early stopping logic to be executed efficiently and consistently across diverse hardware (CPUs, GPUs, custom accelerators) and software environments (TensorFlow, PyTorch, Caffe2) without requiring the entire training stack to be uniform. The communication of the stop signal would occur through standard API calls (e.g., Python function calls) from the ONNX Runtime host to the NN training process.
    • Obviousness: Given the industry-wide focus on model interoperability and efficient deployment, it would be obvious for a PHOSITA to leverage a universal model exchange format like ONNX for the predictive early stopping model. Standardizing the representation and execution of the predictive logic ensures broader applicability and reduces integration complexity across heterogeneous ML ecosystems.
  3. US11650968 + Prometheus Monitoring for Resource-Aware Early Stopping:

    • Description: The predictive early stopping method is augmented with real-time resource utilization monitoring through the open-source Prometheus system. Alongside the target NN's training, custom Prometheus exporters running on the training infrastructure (e.g., GPU servers, compute clusters) continuously scrape and expose metrics such as GPU utilization percentage, VRAM consumption, CPU load, network bandwidth, and instantaneous power draw. This rich set of resource metrics, timestamped and stored in Prometheus's time-series database, is incorporated into an extended "loss history" for the early stopping predictive model. The predictive model (Claim 1), now trained on historical data correlating loss trajectories with resource consumption, determines the probability of achieving further loss reduction with a favorable resource efficiency. Prometheus Alertmanager rules can be configured to trigger a soft stop or a "low-power training mode" if the predicted resource consumption for marginal gain becomes excessively high, even before pure loss-based stagnation. This integration provides a comprehensive, resource-aware early stopping strategy within a standard, widely adopted monitoring framework.
    • Obviousness: For a PHOSITA in cloud computing or MLOps concerned with optimizing infrastructure costs and efficiency, it would be obvious to combine a sophisticated early stopping mechanism with an industry-standard monitoring solution like Prometheus. The ability to collect and analyze granular resource metrics in real-time naturally extends the decision-making capabilities of an early stopping algorithm beyond just performance metrics to include operational efficiency.

Generated 5/16/2026, 6:48:59 PM

Keep exploring

More patents asserted by Unified Patents

Other patents in High-Tech (T)

See all High-Tech (T) patents →

This patent in court (1)

1 tracked lawsuit name US 11650968.