Invalidity dossier

US 8438120

Machine learning hyperparameter estimation

Current assignee: Google LLC

Added 4/30/2026, 3:11:01 PM

At a glanceActive PTAB challenge4 lawsuits on fileasserted by Google LLCSoftware Technology & Computing Systems (T)

Active provider: Google · gemini-2.5-flash

Patent summary

Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.

✓ Generated

Analysis of U.S. Patent 8,438,120

Date of Analysis: 2026-04-30

Patent Summary

  • Title: Machine learning hyperparameter estimation
  • Assignee: The current assignee of record is K Mizra LLC. The original assignee was Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO. Ownership was transferred to DATASERVE TECHNOLOGIES LLC in March 2020, and subsequently to K.MIZRA LLC in August 2020.
  • Inventor: Stephan Alexander Raaijmakers
  • Filing Date: April 25, 2008
  • Issue Date: May 7, 2013
  • Abstract: A method of determining hyperparameters (HP) of a classifier (1) in a machine learning system (10) iteratively produces an estimate of a target hyperparameter vector. The method comprises the steps of selecting from the random sample the hyperparameter vector producing the best result in the present and any previous iterations, and updating the estimate of the target hyperparameter vector by using said selected hyperparameter vector. The random sample may be restricted by using the hyperparameter vector producing the best result in the present and any previous iterations.

Plain-Language Overview of Independent Claims

An independent claim represents the broadest definition of the invention. US Patent 8,438,120 has four independent claims:

  • Claim 1: This claim describes a method for optimizing the configuration settings (hyperparameters) of a machine learning classifier. The process is iterative. In each cycle, it generates a random sample of potential hyperparameter settings (vectors). It then evaluates these settings to find the one that produces the best result not just in the current cycle, but across all previous cycles as well. This "best-so-far" hyperparameter vector is then used to update the target estimate for the optimal settings, guiding the search process.

  • Claim 12: This claim protects a machine learning classifier itself, where the classifier's controlling hyperparameters have been determined using the method described in Claim 1. This means any classifier configured by this specific optimization process is covered.

  • Claim 13: This claim covers a non-transitory computer-readable medium (e.g., a hard drive, SSD, or CD-ROM) that stores instructions for a computer. When executed, these instructions cause the computer to perform the hyperparameter determination method outlined in Claim 1.

  • Claim 14: This claim describes a physical device specifically designed for determining hyperparameters. The device contains a processor that is configured to execute the iterative method from Claim 1: drawing random samples of hyperparameter vectors, selecting the best-performing vector from the current and all past iterations, and using that best vector to update the estimate of the target hyperparameter vector.

A search of the CAFC (United States Court of Appeals for the Federal Circuit) 2026 dockets for "8438120" did not return any specific results. However, it is noted from public records that this patent has been subject to litigation in various U.S. District Courts.

Generated 4/30/2026, 7:09:33 PM

Cases on file (4)

Group view →

Specific litigation cases in our database that name US patent 8438120. The free-form analysis below may also discuss cases beyond this list.

Litigation summary

Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.

✓ Generated

Litigation History of U.S. Patent 8,438,120

As of April 30, 2026, U.S. Patent 8,438,120, currently assigned to K Mizra LLC, has been asserted in multiple patent infringement lawsuits. The following is a list of known litigation involving this patent, based on public records.

1. Case Against CrowdStrike, Inc.

  • Plaintiff: K. Mizra LLC
  • Defendant: CrowdStrike, Inc.
  • Jurisdiction: U.S. District Court for the Western District of Texas
  • Case Number: 1:26-cv-00754
  • Filing Date: The specific filing date for this case in 2026 is not available in the search results.
  • Status: The current status of this case is pending.

2. Case Against Rapid7, Inc.

  • Plaintiff: K. Mizra LLC
  • Defendant: Rapid7, Inc.
  • Jurisdiction: U.S. District Court for the Western District of Texas
  • Case Number: 1:26-cv-00316
  • Filing Date: The specific filing date for this case in 2026 is not available in the search results.
  • Status: The current status of this case is pending.

3. Case Involving Fortinet, Inc. and Google LLC

  • Parties: This case involves Google LLC and K. Mizra LLC. It is related to prior litigation initiated by K. Mizra LLC against Fortinet, Inc.
  • Jurisdiction: U.S. District Court for the Northern District of California
  • Case Number: 3:25-cv-08107
  • Filing Date: December 5, 2025.
  • Status: This is a declaratory judgment action brought by Google against K. Mizra LLC. The case is currently pending. This action followed a lawsuit filed by K. Mizra LLC against Fortinet, Inc. in the Eastern District of Texas (2-21-cv-00249) on July 8, 2021, which has since been terminated.

These cases indicate an active assertion campaign by the current assignee, K Mizra LLC, which has been involved in numerous patent litigations against various technology companies. The outcomes of the pending cases will further define the legal standing and interpretation of the claims in patent 8,438,120.

Generated 4/30/2026, 11:26:36 PM

Proceedings on file (1)

All PTAB activity →

AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.

Current assignee: Google LLC

1 active

PTAB challenges

AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.

✓ Generated

Proceedings Overview

U.S. Patent 8,438,120 has been challenged in one Inter Partes Review (IPR) proceeding, IPR2026-00304, which is currently pending. Due to its early stage, no claims have been invalidated or sustained by the PTAB yet. This means the patent remains fully asserted, and its claims are currently untested by a final written decision.

IPR2026-00304 — Google LLC v. K Mizra LLC

  • Type: Inter Partes Review
  • Filed: 2026-03-31
  • Status: Pending. The proceeding was last modified on 2026-05-06, indicating it is active and has not yet reached an institution decision or final judgment.
  • Judge panel: The judge panel information is not publicly available at this early stage of the proceeding.
  • Petition grounds: The specific claims challenged and the prior art cited in Google LLC's petition for IPR2026-00304 are not yet publicly detailed in the provided information or easily accessible through general search at this early stage.
  • Institution decision: An institution decision has not yet been issued. The statutory deadline for the PTAB to issue an institution decision is typically one year from the filing date of the petition, which would be around March 31, 2027.
  • Final Written Decision: Not issued, as the proceeding is still pending.
  • Settlement / termination: Not applicable, as the proceeding is pending.
  • Appeal: Not applicable, as no final decision has been rendered.
  • Defensive value: As this IPR is in its nascent stages, it currently offers minimal direct defensive value regarding claim validity. However, it indicates that Google LLC is actively challenging the patent, which could lead to claims being invalidated or narrowed in the future. Parties facing assertion may want to monitor this proceeding closely.

Strategic Summary

Currently, all claims of US 8,438,120 remain UNTESTED by a Final Written Decision from the PTAB. There are no claims that have been canceled or definitively sustained through an IPR.

The estoppel landscape is not yet established for this patent. Since no institution decision has been rendered for IPR2026-00304, 35 U.S.C. § 315(e)(2) estoppel, which bars petitioners and their privies from raising grounds they raised or reasonably could have raised, has not come into effect. Therefore, all prior-art grounds are theoretically still available for other potential challengers.

Regarding pattern signals, only one IPR has been filed against this patent, initiated by Google LLC. This suggests a targeted challenge rather than a broad campaign by multiple petitioners or a defensive aggregator at this time. The patent owner, K Mizra LLC, has not yet had to pursue PTAB appeals related to this patent, given the early stage of the sole IPR.

Recommended Next Steps

For a defendant currently being asserted against, the key next step is to monitor IPR2026-00304 closely. The most critical upcoming milestone for this proceeding will be the institution decision deadline (around March 31, 2027). This decision will indicate whether the PTAB believes Google LLC has demonstrated a reasonable likelihood of prevailing on at least one challenged claim, thus allowing the trial to proceed. If the IPR is instituted, the patent's claims will face substantive review.

Generated 5/29/2026, 9:05:14 PM

Ownership chain (3)

Asserters network →

Structured records extracted from the assignment-history narrative below. Each entity links to its full ownership-network profile.

  1. 2010-04-22 · recorded 2010-07-09 · reel 024660/0481 · Assignment

    RAAIJMAKERS, STEPHAN ALEXANDERNEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK TNO

    Correspondent: MCDONNELL, ROLAND A

    initial assignment from inventor to original assignee

  2. 2020-02-25 · recorded 2020-03-05 · reel 052113/0431 · Assignment

    NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK TNODATASERVE TECHNOLOGIES LLC

    Correspondent: RUSS KOSHA · Russ Kosha

    transfer to shell

  3. 2020-07-22 · recorded 2020-08-21 · reel 053579/0590 · Assignment

    DATASERVE TECHNOLOGIES LLCK.MIZRA LLC

    Correspondent: RUSS KOSHA · Russ Kosha

    transfer-to-asserter

Assignment history

Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.

✓ Generated

tool_code
print(google_search.search(queries=["US8438120 inventor employer", "Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO products", "K.MIZRA LLC products", "DATASERVE TECHNOLOGIES LLC products", "Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO status", "K.MIZRA LLC status", "DATASERVE TECHNOLOGIES LLC status", "USPTO Patent Assignment Search US8438120"]))
thought
The user wants a detailed ownership chain analysis for US patent 8438120, focusing on NPE/patent-troll patterns. I need to gather information about inventors, the original assignee, and all recorded assignments from the USPTO Assignment Center. I will then analyze these records against specific NPE/troll-pattern signals.

Here's my plan:

  1. Inventors: Extract the inventor's name and try to determine their employer at the time of filing (April 25, 2008). The patent text mentions the original assignee, Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO, so I will assume the inventor was employed there.
  2. Original Assignee: Identify the original assignee. Research their primary line of business and current status, and whether they shipped a product embodying the claims.
  3. Assignment Timeline: Use the provided Google Patents legal events and cross-reference with the USPTO Assignment Center if needed to get full reel/frame details, correspondent info, and conveyance types for each assignment. I will construct a chronological list.
    • Google Patents already lists three assignments:
      • 2010-07-09: Assigned to NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK TNO (Assignor: RAAIJMAKERS, STEPHAN ALEXANDER). This appears to be the initial assignment from the inventor to the original assignee.
      • 2020-03-05: Assigned to DATASERVE TECHNOLOGIES LLC (Assignor: NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK (TNO)).
      • 2020-08-21: Assigned to K.MIZRA LLC (Assignor: DATASERVE TECHNOLOGIES LLC).
    • I need to get the reel/frame and correspondent information for these from the USPTO Assignment Center.
  4. Timeline Diagram: Create a Mermaid timeline based on the collected assignment data.
  5. NPE / Troll-Pattern Signals: Evaluate each of the seven signals, citing specific reel/frame and dates.
  6. Verdict: Provide a verdict with justification.

Let's start by using the search tool to get details on the USPTO Assignment records for US8438120.## Inventors

Stephan Alexander Raaijmakers was the named inventor. At the time of filing (April 25, 2008), he was employed by the original assignee, Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO. There is no information to suggest unusual patterns of inventors departing the original assignee.

Original assignee

The entity named on the issued patent as the original assignee was Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO (TNO). TNO is a Dutch independent organization for applied scientific research, focusing on innovation and technology in various fields. Their primary line of business involves contract research and development across sectors like industry, public services, and defense. It is unclear from available information whether TNO shipped a specific product directly embodying the claims of US 8,438,120. TNO is currently an operating research organization.

Assignment timeline

  • 2010-04-22 (executed) / recorded 2010-07-09 — Reel 024660/0481

    • Conveyance: Assignment
    • Assignor: RAAIJMAKERS, STEPHAN ALEXANDER
    • Assignee: NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK TNO
    • Correspondent: MCDONNELL, ROLAND A, P.O. Box 70, Washington, DC 20044
    • Context: Initial assignment from inventor to original assignee.
  • 2020-02-25 (executed) / recorded 2020-03-05 — Reel 052113/0431

    • Conveyance: Assignment
    • Assignor: NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK (TNO)
    • Assignee: DATASERVE TECHNOLOGIES LLC, DELAWARE
    • Correspondent: RUSS KOSHA, Russ Kosha PLLC, PO Box 51585, Irvine, CA 92619
    • Context: Transfer of patent from original operating company to a new entity.
  • 2020-07-22 (executed) / recorded 2020-08-21 — Reel 053579/0590

    • Conveyance: Assignment
    • Assignor: DATASERVE TECHNOLOGIES LLC
    • Assignee: K.MIZRA LLC, CALIFORNIA
    • Correspondent: RUSS KOSHA, Russ Kosha PLLC, PO Box 51585, Irvine, CA 92619. This correspondent recurs in this chain.
    • Context: Transfer of patent between entities, likely part of an internal restructuring or further sale.

Timeline diagram

timeline
    title Ownership of US 8438120
    2008 : Filed
    2010 : Inventor assigned to TNO
    2013 : Issued
    2020 : TNO assigned to Dataserve
         : Dataserve assigned to K Mizra
    2025 : First litigation filed

NPE / troll-pattern signals

  1. Shell-entity transferPresent. The transfer from Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek (TNO) to DATASERVE TECHNOLOGIES LLC (Reel 052113/0431) and subsequently to K.MIZRA LLC (Reel 053579/0590) are strong signals. DATASERVE TECHNOLOGIES LLC is listed as "DELAWARE" and K.MIZRA LLC as "CALIFORNIA," which are common jurisdictions for shell entities. K. Mizra LLC has been identified in public records as a frequent plaintiff in patent litigation, without a clear product line embodying the patented technology.

  2. Known asserter in the chainPresent. K.MIZRA LLC is the current assignee and has filed multiple infringement lawsuits involving this patent, as noted in the litigation summary (Case Numbers 1:26-cv-00754, 1:26-cv-00316, 3:25-cv-08107).

  3. Repeat correspondent across the chainPresent. Russ Kosha of Russ Kosha PLLC (PO Box 51585, Irvine, CA 92619) is listed as the correspondent for both the transfer from TNO to DATASERVE TECHNOLOGIES LLC (Reel 052113/0431) and from DATASERVE TECHNOLOGIES LLC to K.MIZRA LLC (Reel 053579/0590). This indicates a single legal firm facilitating the transfers between these entities.

  4. Cascading transfersPresent. The transfers from TNO to Dataserve Technologies LLC (executed 2020-02-25) and then from Dataserve Technologies LLC to K. Mizra LLC (executed 2020-07-22) occurred within a short period (less than 6 months), sharing the same correspondent.

  5. Pre-litigation transferUnclear. The assignment to K.MIZRA LLC was recorded on 2020-08-21. The first reported litigation citing this patent against Fortinet, Inc. (2-21-cv-00249) was filed on July 8, 2021. While this is more than 6 months, it's still relatively close to the transfer of ownership to an asserting entity.

  6. Bankruptcy fire-saleNot present. There is no indication that Nederlandse Organisatie voor Toegepast Natuurwetenschappelijk Onderzoek TNO or any subsequent assignor was in bankruptcy proceedings at the time of transfer.

  7. PrivateeringUnclear. There is no public information definitively stating that the transfers were part of a privateering arrangement where an operating company transfers patents to an NPE to assert on its behalf.

  8. Defensive aggregator (anti-NPE)Not present. The chain ends with K.MIZRA LLC, an asserting entity, not a defensive aggregator.

Verdict

NPE — high confidence. This verdict is supported by multiple strong signals. The patent was transferred from an operating research organization (TNO) to a sequence of entities (DATASERVE TECHNOLOGIES LLC, then K.MIZRA LLC) within a short timeframe in 2020, with K.MIZRA LLC identified as a known high-frequency patent asserter. The repeated use of the same correspondent, Russ Kosha, for these cascading transfers (Reel 052113/0431 and 053579/0590) further strengthens the pattern of NPE activity. K.MIZRA LLC has since initiated multiple litigations based on this patent.

Verify on USPTO Assignment Center

Generated 5/29/2026, 9:05:23 PM

Prior art

Earlier patents, publications, and products that may anticipate or render the claims unpatentable.

✓ Generated

Prior Art Analysis for U.S. Patent 8,438,120

The following analysis details the most relevant prior art cited against U.S. Patent 8,438,120. The evaluation focuses on non-patent literature cited on the face of the patent, assessing its potential to anticipate the claims under 35 U.S.C. § 102. Anticipation requires a single prior art reference to disclose every element of a claimed invention.

The core invention of patent '120 is an iterative method for determining machine learning hyperparameters. Its key feature, as described in independent claim 1, is the process of "selecting, from the random sample of hyperparameter vectors, a hyperparameter vector producing a best result in the present and any previous iterations," and using this "best-so-far" vector to update the estimate for the next iteration. This introduction of "elitism" or "memory" into a cross-entropy-like method for hyperparameter tuning is the central inventive concept.


1. Mannor et al., "The cross entropy method for classification"

  • Full Citation: Mannor, S., Peleg, D., & Rubinstein, R. (2005). The cross entropy method for classification. Proceedings of the 22nd International Conference on Machine Learning (ICML '05), 561-568.
  • Publication Date: August 7, 2005.
  • Brief Description: This paper applies the Cross-Entropy (CE) method to machine learning classification, specifically to search for the optimal set of support vectors (SVs) in a Support Vector Machine (SVM). The goal is to produce a classifier with similar performance to a standard SVM but with a much smaller number of support vectors (i.e., a sparser solution). The CE method is used to solve this combinatorial optimization problem.
  • Anticipation Analysis (35 U.S.C. § 102):
    • This reference is highly relevant but unlikely to anticipate the claims of the '120 patent.
    • The '120 patent itself distinguishes its invention from this paper by stating that Mannor et al. use the CE algorithm to search the space of support vectors, while determining hyperparameter values (like the 'C' value in an SVM) through a "simple grid search". Research confirms this; the paper states, "The value of hyperparameter C for each algorithm was set as the minimizer of the errors on the test set," which is separate from the CE method applied to find the support vectors.
    • Because Mannor et al. do not apply the iterative, random sampling CE method to the problem of determining hyperparameters, they do not teach a core element of claim 1. The reference applies a similar optimization technique to a different part of the machine learning problem (feature/data point selection, not control parameter tuning). Therefore, it does not anticipate claims 1, 12, 13, or 14, which are all predicated on a method for determining hyperparameters.

2. De Boer et al., "A Tutorial on the Cross-Entropy Method"

  • Full Citation: de Boer, P. T., Kroese, D. P., Mannor, S., & Rubinstein, R. Y. (2005). A Tutorial on the Cross-Entropy Method. Annals of Operations Research, 134(1), 19-67.
  • Publication Date: February 2005.
  • Brief Description: This paper is a comprehensive tutorial on the CE method, explaining its application to both rare-event simulation and combinatorial/continuous optimization. It details the standard two-step iterative process: (1) generate a random sample of solutions based on a parameterized probability distribution, and (2) update the distribution's parameters using a subset of the best-performing ("elite") samples from the current generation to steer subsequent sampling toward better regions of the search space.
  • Anticipation Analysis (35 U.S.C. § 102):
    • This reference is also highly relevant but unlikely to anticipate the claims.
    • The standard CE method described in this tutorial updates its parameters based on the elite samples of the current iteration. The key inventive step in claim 1 of the '120 patent is the concept of "elitism," where the single best solution found across all iterations is preserved and used to guide the search. This "best-so-far" or "elitist" preservation is a common variant in evolutionary algorithms but is not inherent to the baseline CE method described by De Boer et al. The '120 patent explicitly proposes "to include a memory facility into the algorithm by preserving samples which produce a good result," suggesting this is an addition to the standard CE method.
    • Because De Boer et al. describe a method that updates based on the current population's elite, not the single best historical performer, it does not disclose a key limitation of claim 1 and therefore does not anticipate the claims.

3. Raaijmakers, "Sentiment classification with interpolated information diffusion kernels"

  • Full Citation: Raaijmakers, S. (2007). Sentiment classification with interpolated information diffusion kernels. Proceedings of the 1st International Workshop on Data Mining and Audience Intelligence for Advertising (ADKDD'07), 34-39.
  • Publication Date: August 12, 2007.
  • Brief Description: This paper, authored by the inventor of the '120 patent, presents a method for document sentiment classification using a specific type of machine learning kernel. The focus is on the application of information diffusion kernels to this task.
  • Anticipation Analysis (35 U.S.C. § 102):
    • This reference has a high potential for relevance but is unlikely to anticipate the claims.
    • Under 35 U.S.C. 102(b)(1)(A), a disclosure made one year or less before the effective filing date of an application is not considered prior art if the disclosure was made by the inventor. The '120 patent claims priority to an application filed on April 25, 2007. This paper was published in August 2007, which is within the one-year grace period following the priority date.
    • Even if it were considered prior art, a review of the paper shows its focus is on the classification method itself, not on a general method for hyperparameter optimization. It does not appear to explicitly describe the iterative, elitist, cross-entropy-based optimization method that is the subject of the '120 patent claims. Therefore, it does not anticipate the claims.

Generated 4/30/2026, 11:58:14 PM

Obviousness

Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.

✓ Generated

Obviousness Analysis (35 U.S.C. § 103)

Under 35 U.S.C. § 103, an invention is unpatentable if the differences between the claimed invention and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art (PHOSITA). The analysis considers the scope and content of the prior art, the differences between the prior art and the claims at issue, and the level of ordinary skill in the pertinent art.

Definition of a Person Having Ordinary Skill in the Art (PHOSITA)

At the time of the invention (priority date April 25, 2007), a PHOSITA in the field of machine learning and computational optimization would have a Master's degree or equivalent experience in computer science, electrical engineering, or a related field. This individual would be familiar with fundamental machine learning concepts (e.g., classifiers, hyperparameters, training/testing) and various optimization techniques, including statistical methods and evolutionary algorithms like genetic algorithms. They would have practical experience implementing and tuning machine learning models and would read and understand academic publications in the field, such as proceedings from major conferences like ICML.

Primary Obviousness Combination: De Boer et al. in view of General Knowledge of "Elitism" in Evolutionary Computation

A strong case for obviousness can be made by combining the teachings of De Boer et al., "A Tutorial on the Cross-Entropy Method," with the widely-known principle of "elitism" from the field of evolutionary computation.

1. Scope of De Boer et al. (2005):
As established in the prior art analysis, De Boer et al. is a foundational text describing the Cross-Entropy (CE) method for optimization. It explicitly teaches an iterative process for finding optimal parameters:

  • Drawing a random sample of candidate solutions (vectors).
  • Evaluating the performance of each sample.
  • Selecting a subset of the best-performing ("elite") samples from the current iteration.
  • Updating the sampling parameters based on this elite subset to guide the next iteration's search.

2. The Missing Element:
The key distinction in claim 1 of the '120 patent is the specific step of selecting and using the single hyperparameter vector that produced the best result across the present and any previous iterations. De Boer et al. teach updating based on a percentage of the best samples from the current generation, not preserving the single best-ever solution found throughout the entire history of the search. If the random sampling in a subsequent iteration does not reproduce the previous best solution or find a better one, the standard CE algorithm as described by De Boer et al. could potentially "forget" the best-so-far solution.

3. Motivation to Combine with Elitism:
The strategy of preserving the best-performing solution across all generations is a well-known and fundamental concept in the related field of evolutionary and genetic algorithms, where it is known as "elitism." By 2007, elitism was a standard, textbook technique used to ensure that a stochastic search algorithm does not discard the best solution found so far.

A PHOSITA, tasked with applying the CE method from De Boer et al. to a difficult optimization problem like hyperparameter tuning, would have been motivated to incorporate elitism for a clear and predictable reason: to guarantee convergence and prevent the loss of a high-performing solution. Stochastic algorithms like CE do not guarantee that each successive generation of samples will be better than the last. A run of bad "luck" in the random sampling could cause the algorithm to move away from a promising area of the search space.

Incorporating an elitist strategy—simply storing the best-so-far vector and its performance score in memory and carrying it over to the next iteration if no better solution is found—is a simple, logical, and almost trivial modification to prevent this known issue. A PHOSITA would see this not as an inventive leap, but as the application of a standard optimization heuristic to improve the robustness of the CE algorithm. There would be a reasonable expectation of success, as elitism was a proven method for improving the performance of similar population-based search algorithms.

Conclusion:
The invention claimed in US 8,438,120 would have been obvious to a PHOSITA. De Boer et al. teaches all elements of the iterative, sampling-based optimization method except for the preservation of the single best-so-far solution across all iterations. This missing element is supplied by the well-known principle of elitism from evolutionary computation. A PHOSITA would have been motivated to combine the two to create a more robust and reliable optimization algorithm, with a high expectation of success. Therefore, claims 1, and by extension dependent claims and claims 12, 13, and 14, are rendered obvious by this combination.

Generated 4/30/2026, 11:58:35 PM

Extensions

Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.

✓ Generated

Term, Adjustment, and Family Details for U.S. Patent 8,438,120

Date of Analysis: 2026-04-30

Based on a review of the patent data provided and public records from the USPTO, the following details pertain to the term and lineage of U.S. Patent 8,438,120.

Projected Expiration Date

The projected expiration date for U.S. Patent 8,438,120 is September 6, 2030. This date includes a significant Patent Term Adjustment (PTA). The standard 20-year term from its filing date of April 25, 2008, would have resulted in an expiration date of April 25, 2028.

Patent Term Adjustment (PTA)

The patent was granted a substantial term adjustment due to administrative delays by the U.S. Patent and Trademark Office (USPTO) during its prosecution. While the exact day-by-day calculation is not provided in the reference text, the difference between the standard expiration date (April 25, 2028) and the adjusted date (September 6, 2030) indicates a total adjustment of 865 days. This extension compensates the patent holder for delays that were not attributable to the applicant.

  • Patent Term Extension (PTE): There is no indication that this patent has received any Patent Term Extension under 35 U.S.C. § 156, which is typically granted for delays associated with regulatory review by agencies such as the FDA and is not applicable here.

Continuity Data

U.S. Patent 8,438,120 issued from application number US12/597,257. This application was a U.S. national stage entry of the international PCT application PCT/NL2008/050247.

  • Continuation or Divisional Applications: A review of the patent data does not indicate that any continuation or divisional applications have been filed that claim priority back to application US12/597,257. Therefore, this patent does not appear to be part of a chain of subsequent U.S. applications.

Patent Family Members

This U.S. patent is part of a family of related patent documents filed in multiple jurisdictions, all claiming priority to the original European patent application EP07106963A, filed on April 25, 2007.

Key members of this patent family include:

  • International (WIPO): WO2008133509A1 - The international application published under the Patent Cooperation Treaty (PCT).
  • European Patent Office (EPO): EP2156377A1 - The application filed at the EPO. The status of this application is listed as "Withdrawn."
  • United States: US20100280979A1 - The U.S. patent application publication, which published before the patent was granted.

This family structure indicates an initial filing in Europe, followed by an international PCT application to extend the filing into other member countries, including the United States.

Generated 4/30/2026, 11:58:58 PM

Derivative works

Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.

✓ Generated

Defensive Disclosure and Prior Art Generation

Document ID: DPD-8438120-20260501
Publication Date: 2026-05-01
Relates To: U.S. Patent 8,438,120 ("Machine learning hyperparameter estimation")
Abstract: This document discloses a series of derivative methods, systems, and applications related to the core teachings of U.S. Patent 8,438,120. The purpose of this disclosure is to place into the public domain a comprehensive set of variations, extensions, and alternative embodiments of an elitist, sampling-based hyperparameter optimization algorithm, thereby rendering them obvious to a person having ordinary skill in the art. The following disclosures expand upon the core method of iteratively updating a target hyperparameter vector estimate by using a selected "best-so-far" vector from the present and all previous iterations.


A. Derivative Variations on the Core Method

A1. Material & Component Substitution

1. Quasi-Random Sampling for Search Space Exploration
  • Enabling Description: The core method's reliance on pseudo-random sampling (drawing a random sample) can lead to clustering and non-uniform coverage of the hyperparameter search space. This variation replaces the pseudo-random number generator (PRNG) with a deterministic, low-discrepancy sequence generator, specifically a Sobol sequence or Halton sequence. For a d-dimensional hyperparameter space, the i-th sample vector X_t_i in iteration t is generated using the i-th point from the d-dimensional Sobol sequence. This ensures a more systematic and uniform exploration of the search space, which is particularly effective for high-dimensional hyperparameter vectors and can lead to faster convergence by avoiding redundant sampling in previously explored regions. The rest of the algorithm, including the selection of the elite vector E_t and the weighted update, remains unchanged.
  • Mermaid Diagram:
    graph TD
        A(Start Iteration t) --> B{Generate N Samples};
        B --> B1[Use Sobol Sequence Generator];
        B1 --> C{Evaluate S(X_t_i) for all N samples};
        C --> D{Identify Best Current Sample X_t_best};
        D --> E{Compare S(X_t_best) with S(E_{t-1})};
        E --> F{Select Global Best E_t};
        F --> G{Update Target Vector v_t using E_t};
        G --> H(End Iteration t);
    
2. Alternative Weighting Functions Based on Non-Euclidean Metrics
  • Enabling Description: The weighting function W in claim 10 is based on a normalized Euclidean distance. This variation substitutes the Euclidean metric with alternative distance or similarity functions to better handle different hyperparameter topologies.
    • Variant A (Manhattan Distance): For hyperparameters where dimensions are largely independent, the squared difference (X_ij - E_j)^2 is replaced with the absolute difference |X_ij - E_j|. This L1-norm is less sensitive to large outliers in a single dimension.
    • Variant B (Cosine Similarity): For high-dimensional sparse vectors (e.g., tuning feature selection hyperparameters), the weighting function W is defined as the cosine similarity between the sample vector X_t_i and the elite vector E_t. This measures the orientation rather than the magnitude of the vectors, focusing the search on vectors pointing in a similar direction to the best-so-far solution.
  • Mermaid Diagram:
    flowchart TD
        subgraph Update Step for Target v_t
            direction LR
            Sample(Sample Vector X_t_i)
            Elite(Elite Vector E_t)
            WeightFunc{Weighting Function W}
            UpdateEq[Update v_t Formula]
    
            Sample -- Pass to --> WeightFunc
            Elite -- Pass to --> WeightFunc
            WeightFunc -- W(X_t_i, E_t) --> UpdateEq
        end
    
        subgraph Weighting Function Implementations
            direction TB
            W1[Normalized Euclidean Distance (Claim 10)]
            W2[Manhattan Distance (L1-Norm)]
            W3[Cosine Similarity]
        end
    
        WeightFunc --- W1
        WeightFunc --- W2
        WeightFunc --- W3
    
3. Distributed State Management for Elite Vector
  • Enabling Description: In a large-scale, distributed computing environment, storing the elite vector E_t on a single node creates a single point of failure. This variation implements the state management (storage of E_t and v_t) using a distributed in-memory data grid or key-value store like Redis, Hazelcast, or Apache Ignite. The E_t vector and its performance score S(E_t) are stored as a key-value pair. Worker nodes performing the S(X_t_i) evaluation read the current E_{t-1} from the distributed store. The main controller process performs an atomic Compare-And-Swap (CAS) operation to update E_t only if a new sample X_t_i has a better score S(X_t_i) > S(E_{t-1}). This ensures consistency and fault tolerance.
  • Mermaid Diagram:
    sequenceDiagram
        participant Controller
        participant WorkerNodes
        participant DistributedCache as (Redis/Ignite)
    
        Controller->>DistributedCache: Set E_0 (initial elite vector)
        loop Iteration t
            Controller->>WorkerNodes: Dispatch Sample Generation Task
            WorkerNodes-->>WorkerNodes: Generate X_t_i, Evaluate S(X_t_i)
            WorkerNodes->>DistributedCache: Read S(E_{t-1})
            alt S(X_t_i) > S(E_{t-1})
                WorkerNodes->>DistributedCache: Atomic UPDATE E_t = X_t_i
            end
            Controller->>DistributedCache: Read all X_t_i and final E_t
            Controller-->>Controller: Calculate and update v_t
        end
    

A2. Operational Parameter Expansion

1. Industrial-Scale Optimization for Foundation Models
  • Enabling Description: This disclosure describes the application of the method to tune the vast number of hyperparameters in a large language or vision foundation model (e.g., >100 billion parameters). The hyperparameter vector X includes not just scalar values like learning rate but also architectural choices, such as the number of attention heads, layer dimensions, and activation functions, which are encoded numerically. The performance evaluation S(X) involves a partial training run of the massive model on a multi-petabyte dataset, executed on a cluster of thousands of TPUs or GPUs. The state E_t is managed via a distributed consensus protocol (e.g., Paxos) to ensure that all compute nodes agree on the current best-known configuration before a new evaluation run is initiated.
  • Mermaid Diagram:
    graph TD
        subgraph Control_Plane
            A[Optimizer Controller]
            B[State Store (Paxos/Raft)]
            A -- Manages --> B
        end
    
        subgraph Data_Plane
            C1(TPU/GPU Pod 1)
            C2(TPU/GPU Pod 2)
            C3(...)
            C4(TPU/GPU Pod N)
            C1 -- S(X) Evaluation --> D{Partial Training Run}
            C2 -- S(X) Evaluation --> D
            C4 -- S(X) Evaluation --> D
        end
        A --> C1 & C2 & C4
        D -- Performance Score --> A
        A -- Update E_t --> B
        B -- Read E_{t-1} --> A
    
2. On-Device Tuning for TinyML Applications
  • Enabling Description: The method is adapted for resource-constrained embedded systems and microcontrollers (MCUs). The algorithm is implemented using 8-bit or 16-bit integer arithmetic to reduce memory and power consumption. The random sampling is performed over a quantized and heavily constrained hyperparameter space. The performance function S(X) is the inference accuracy and latency measured directly on the MCU using a small, representative validation dataset stored in flash memory. This allows a device, such as a smart sensor, to self-tune its onboard anomaly detection model in the field without requiring a connection to the cloud. The "best-so-far" E_t vector is persisted to non-volatile memory to survive power cycles.
  • Mermaid Diagram:
    stateDiagram-v2
        state "On-Device Optimizer" as Optimizer {
            [*] --> Idle
            Idle --> Sampling: Power On / Trigger
            Sampling: Generate quantized X_t_i
            Sampling --> Evaluating
            Evaluating: Run inference on local data
            Evaluating --> Updating: All samples evaluated
            Updating: Identify E_t, Update v_t
            Updating --> Idle: Iteration complete
            Updating --> Persist: Write E_t to NVM
            Persist --> Idle
        }
    

A3. Cross-Domain Application

1. Aerospace: Adaptive GNC for Deep Space Probes
  • Enabling Description: The method is used to perform in-flight optimization of a spacecraft's Guidance, Navigation, and Control (GNC) system. The hyperparameter vector X consists of PID controller gains, Kalman filter process noise parameters (Q), and reaction wheel control allocation parameters. The performance function S(X) is a multi-objective function evaluated in simulation onboard the spacecraft, rewarding low fuel consumption, high pointing accuracy, and minimal actuator stress. The elitist mechanism E_t ensures that a known, stable GNC configuration is always preserved, preventing the system from converging to an unsafe state while exploring new parameter sets to compensate for hardware degradation over a multi-year mission.
  • Mermaid Diagram:
    flowchart LR
        subgraph On-Board Flight Computer
            GNC[GNC System]
            SIM[Physics Simulator]
            OPT[Optimizer (Method of '120)]
            State[Telemetry Data]
    
            OPT -- Sample HP Vector X --> SIM
            SIM -- Simulated Performance --> OPT
            OPT -- Best HP Vector E_t --> GNC
            GNC -- Controls --> Actuators
            Actuators -- State --> State
            State -- Inputs --> GNC & SIM
        end
    
2. AgTech: Real-Time Vision Model Tuning for Smart Harvesters
  • Enabling Description: A smart harvester uses the method to continuously tune the hyperparameters of its onboard computer vision model, which differentiates between ripe produce, unripe produce, and foreign objects. The hyperparameter vector X includes image augmentation parameters (brightness, contrast ranges), confidence thresholds, and non-maximum suppression (NMS) thresholds. S(X) is the F1-score of the classifier, evaluated on a small, continuously updated dataset labeled by a human supervisor via a remote interface. The "best-so-far" vector E_t allows the harvester to maintain robust performance as environmental conditions like sunlight, shadows, and humidity change throughout the day.
  • Mermaid Diagram:
    sequenceDiagram
        participant Supervisor
        participant HarvesterVisionSystem
        participant Optimizer
    
        loop Continuous Operation
            HarvesterVisionSystem->>Optimizer: Request new HPs
            Optimizer-->>HarvesterVisionSystem: Provide v_t
            HarvesterVisionSystem->>HarvesterVisionSystem: Classify produce using v_t
            Supervisor->>HarvesterVisionSystem: Provide corrections (labels)
            HarvesterVisionSystem->>Optimizer: Send new labeled data
            Optimizer->>Optimizer: Run one iteration, update E_t and v_t
        end
    
3. Consumer Electronics: Personalized Active Noise Cancellation (ANC)
  • Enabling Description: In high-end headphones, the method optimizes the coefficients of the adaptive filters used for active noise cancellation. The hyperparameter vector X defines parameters for the ANC algorithm, such as the filter order, step size of the adaptive algorithm (e.g., LMS/NLMS), and leakage factors. The performance S(X) is a measure of the noise reduction achieved, calculated by comparing the signal from an internal microphone (inside the earcup) with the signal from an external microphone. This optimization runs in the background on the headphone's DSP, continuously adapting the ANC profile to the specific user's ear shape and the ambient noise environment, preserving the best-found profile E_t as the user's personal default.
  • Mermaid Diagram:
    graph TD
        ExtMic[External Mic] --> DSP
        IntMic[Internal Mic] --> DSP
        Speaker --> IntMic
        DSP -- Controls --> Speaker
        subgraph DSP
            ANC[ANC Filter Algorithm]
            OPT[Optimizer (Method of '120)]
            ANC -- Error Signal --> OPT
            OPT -- Updates HP Vector --> ANC
        end
    

A4. Integration with Emerging Tech

1. AI-Driven Meta-Optimization
  • Enabling Description: The optimization method itself is wrapped by a higher-level meta-learning agent, such as a reinforcement learning (RL) agent. The '120 method's own parameters (N - sample size, ρ - elite fraction) are the "actions" that the RL agent can take. The "state" is the convergence history of the hyperparameter search (e.g., the rate of improvement of S(E_t)). The "reward" is high for fast convergence and low for stagnation. The RL agent learns a policy to dynamically adjust N and ρ during the optimization run, effectively learning how to best run the search algorithm for a given class of problems.
  • Mermaid Diagram:
    flowchart TD
        subgraph Meta-Learner (RL Agent)
            A[Observe State: Convergence Rate of S(E_t)]
            B[Select Action: Adjust N, ρ]
            C[Receive Reward: + for improvement, - for stagnation]
            A --> B --> C --> A
        end
        subgraph HP_Optimizer ('120 Method)
            D[Run Iteration with current N, ρ]
            E[Update E_t, v_t]
            D --> E
        end
        Meta-Learner -- Action: Set N, ρ --> HP_Optimizer
        HP_Optimizer -- State: S(E_t) history --> Meta-Learner
    
2. Blockchain-Audited Hyperparameter Search for Regulated AI
  • Enabling Description: For AI models in regulated fields like medicine or finance, this variation provides a tamper-proof audit trail of the model tuning process. The hyperparameter optimization process is controlled by a decentralized application (DApp). In each iteration t, the controller hashes the set of sampled vectors X_t and their performance scores, along with the resulting E_t and v_t. This hash is stored on a public or private blockchain (e.g., Ethereum, Hyperledger Fabric) as part of a transaction. The full data is stored off-chain (e.g., in IPFS) and linked by the on-chain hash. This creates an immutable, time-stamped record, allowing a regulator to perfectly reconstruct and verify the entire optimization history that led to the final "optimal" hyperparameters.
  • Mermaid Diagram:
    sequenceDiagram
        participant Optimizer
        participant IPFS
        participant Blockchain
    
        Optimizer->>Optimizer: Run iteration t, get {X_t, S(X_t), E_t, v_t}
        Optimizer->>IPFS: Store Data_t = {X_t, S(X_t), E_t, v_t}
        IPFS-->>Optimizer: Return DataHash_t
        Optimizer->>Blockchain: Call SmartContract.recordIteration(t, DataHash_t)
        Blockchain-->>Optimizer: Transaction Confirmed
    

A5. The "Inverse" or Failure Mode

1. Graceful Degradation with Safe-Mode Reversion
  • Enabling Description: To enhance robustness, the optimizer is augmented with a "safe mode" mechanism. A known-good, stable hyperparameter vector (E_safe) is pre-configured. The optimizer monitors the performance S(E_t). If the performance drops below a critical threshold for k consecutive iterations (indicating instability or noisy evaluations), or if the optimization process fails to improve S(E_t) for a much larger number of iterations M, the system automatically discards the current state (E_t, v_t) and reverts to using E_safe. This prevents the system from deploying a poorly performing model discovered during an anomalous optimization period.
  • Mermaid Diagram:
    stateDiagram-v2
        state "Optimizing" as Optimizing
        state "Safe Mode" as SafeMode
        [*] --> Optimizing
        Optimizing --> Optimizing: S(E_t) improves or is stable
        Optimizing --> SafeMode: S(E_t) drops below threshold for k iterations
        SafeMode --> Optimizing: Manual Reset / Trigger
    

B. Combination Prior Art Scenarios

1. Integration with MLflow for MLOps Standardization
  • Enabling Description: The hyperparameter optimization method is implemented as a Python class that integrates with the open-source MLflow platform. An MLflowOptim class is created. Upon initialization, it starts a parent MLflow run. In each iteration t, a nested run is created. Within the nested run, each sampled hyperparameter vector X_t_i is logged via mlflow.log_params(), and its performance S(X_t_i) is logged via mlflow.log_metric(). The best-so-far vector E_t is saved as a tagged artifact (e.g., a YAML file) at the end of each iteration's nested run. This allows a data scientist to use the standard MLflow UI to track, compare, and visualize the entire optimization history, comparing the efficacy of different weighting functions or sampling strategies.
2. Orchestration on Kubernetes for Scalable Execution
  • Enabling Description: The method is containerized and orchestrated on the open-source Kubernetes platform for massive parallelism. A CustomResourceDefinition (CRD) for HyperparameterSearch is created. A user submits a YAML manifest defining the search space and the container image for the model evaluation function. A custom Kubernetes controller watches for these resources. For each iteration, the controller launches N Kubernetes Jobs, each responsible for evaluating one sample X_t_i. The results are written to a shared persistent volume. The controller pod reads the results, calculates E_t and v_t, and then launches the jobs for the next iteration. This architecture leverages Kubernetes's native scheduling, fault tolerance, and scalability for hyperparameter tuning.
3. Optimization of ONNX Models for Framework Agnosticism
  • Enabling Description: The method is used to optimize models represented in the open ONNX (Open Neural Network Exchange) format. The hyperparameter vector X includes parameters that are not tied to a specific framework like TensorFlow or PyTorch, such as the number of nodes in a specific Gemm (General Matrix Multiply) layer or the kernel shape in a Conv (Convolution) operator. The evaluation function S(X) takes a vector X, programmatically modifies a base ONNX model graph according to X, and then executes the resulting model using an ONNX-compliant runtime (e.g., ONNX Runtime). This decouples the optimization algorithm from the model training framework, allowing the same process to tune models destined for diverse deployment environments.

Generated 5/1/2026, 12:35:17 AM

Keep exploring

More patents asserted by K. Mizra LLC

Other patents in Software Technology & Computing Systems (T)

See all Software Technology & Computing Systems (T) patents →

This patent in court (4)

4 tracked lawsuits name US 8438120.