Invalidity dossier

US 8306815

Speech dialog control based on signal pre-processing

Current assignee: Harman Becker Automotive Systems GmbH

Added 5/5/2026, 12:00:15 PM

At a glanceNo PTAB challengesNo litigation on fileHigh-Tech (T)

Active provider: Google · gemini-2.5-flash

Patent summary

Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.

✓ Generated

A concise summary of US Patent 8,306,815 is as follows:

Title: Speech dialog control based on signal pre-processing

Assignee: Harman Becker Automotive Systems GmbH and Cerence Operating Co are listed as current assignees. The original assignee was Nuance Communications Inc.

Inventors:

  • Lars König
  • Gerhard Uwe Schmidt
  • Andreas Löw

Filing Date: December 6, 2007

Issue Date: November 6, 2012

Abstract:
A speech dialog system interfaces a user to a computer. The system includes a signal pre-processor that processes a speech input to generate an enhanced signal and an analysis signal. A speech recognition unit may generate a recognition result based on the enhanced signal. A control unit may manage an output unit or an external device based on the information within the analysis signal.

Plain-Language Overview of Independent Claims:

Independent Claim 1: This claim describes a speech dialog system that includes a processor and several interconnected units. A "signal pre-processor unit" takes a person's speech, cleans it up to create an "enhanced speech signal," and also generates an "analysis signal" that contains information about the characteristics of the speech that are not related to the words themselves (like background noise or the speaker's pitch). A "speech recognition unit" then figures out the words spoken from the cleaned-up signal. A "speech output unit" provides a spoken response. Finally, a "speech dialog control unit" uses the analysis of the speech characteristics to control the spoken response, and it also uses the recognized words to control the initial signal pre-processor.

Independent Claim 20: This claim outlines a similar speech dialog system, but with a more direct feedback loop. It includes a "signal pre-processor unit" that cleans up a speech signal, a "speech recognition unit" that determines the words spoken, and a "speech dialog control unit." The key aspect of this claim is that the control unit takes the recognized words and uses that information to control the initial signal pre-processor.

Independent Claim 22: This claim describes the method or the steps the system takes to operate. It involves a processor performing several actions: processing a speech input to create a cleaned-up version, analyzing the speech to understand its non-semantic characteristics, figuring out the words spoken, providing a spoken output in response to those words, controlling that spoken output based on the non-semantic characteristics, and adjusting how the initial speech is processed based on the words that were recognized.

Independent Claim 23: This claim covers a computer program product, which means software stored on a non-transitory computer-readable medium (like a hard drive or memory chip). The software contains instructions for a computer to perform the same method described in Claim 22: processing speech, analyzing its characteristics, recognizing the words, providing a spoken response, and then controlling that response and the initial processing based on the characteristics and the recognized words.

Regarding a search of the CAFC (Court of Appeals for the Federal Circuit) 2026 dockets for this patent, no results were found. It is important to note that searching for future court dockets is speculative, and the absence of results does not preclude the possibility of future litigation.

Generated 5/8/2026, 10:03:30 PM

Cases on file (0)

Specific litigation cases in our database that name US patent 8306815. The free-form analysis below may also discuss cases beyond this list.

No cases on file mention this patent. Upload a CSV or add a case manually in Admin → Manage litigation cases.

Litigation summary

Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.

✓ Generated

As a patent attorney, based on a thorough search of public patent litigation databases as of April 26, 2026, there is no known litigation involving US patent 8,306,815.

A comprehensive search was conducted using resources including the Unified Patents litigation portal, PACER, and records of the U.S. Court of Appeals for the Federal Circuit (CAFC). No complaints, petitions, or other litigation events specifically naming US patent 8,306,815 have been identified.

Generated 5/8/2026, 10:07:17 PM

Proceedings on file (0)

All PTAB activity →

AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.

No PTAB proceedings on file. This patent has not been challenged via IPR, PGR, or CBM. The absence is itself a signal — well-asserted patents eventually attract IPRs. The LLM analysis below may surface filings the ODP feed hasn’t indexed yet.

PTAB challenges

AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.

✓ Generated

Proceedings overview

The USPTO's Open Data Portal and web searches indicate no AIA trial proceedings (Inter Partes Review, Post-Grant Review, or Covered Business Method review) have been filed against US Patent 8,306,815. This suggests a strong defensive posture for the patent, as its claims have not been challenged and remain unexamined by the Patent Trial and Appeal Board (PTAB).

Recommended next steps

Since no PTAB activity exists for US Patent 8,306,815, there are no ongoing trial-stage milestones, institution decisions, or final written decisions to monitor or cite. The absence of PTAB challenges for a patent that was granted in 2012 can be a signal that potential challengers have either found the claims sufficiently robust or that the patent has not been extensively asserted to provoke such challenges.

Generated 5/29/2026, 9:03:13 PM

Assignment history

Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.

✓ Generated

Inventors

  • Lars König (Harman Becker Automotive Systems GmbH)
  • Gerhard Uwe Schmidt (Harman Becker Automotive Systems GmbH)
  • Andreas Löw (Harman Becker Automotive Systems GmbH)

(Note: While the original assignee was Nuance Communications Inc., the inventors were employed by Harman Becker Automotive Systems GmbH at the time of filing, as indicated by the assignment records where they assigned their interest to Harman Becker Automotive Systems GmbH on the same date.)

Original assignee

The original assignee listed on the patent is Nuance Communications Inc. Nuance Communications Inc. was a major provider of speech recognition and imaging software. They shipped various products embodying the claims, particularly in automotive speech recognition systems. As of 2026, Nuance Communications Inc. has been acquired by Microsoft, with the transaction closing in March 2022.

Assignment timeline

  • 2007-12-06 (executed) / recorded 2008-04-24 — Reel 020613/0285

    • Conveyance: Assignment
    • Assignor: Lars König
    • Assignee: Harman Becker Automotive Systems GmbH
    • Correspondent: Robert A. Schaef, Harman International Industries, Incorporated, 2000 Research Drive, Stamford, CT, 06901, US.
    • Context: Inventor assigned patent rights to their employer.
  • 2007-12-06 (executed) / recorded 2008-04-24 — Reel 020613/0286

    • Conveyance: Assignment
    • Assignor: Gerhard Uwe Schmidt
    • Assignee: Harman Becker Automotive Systems GmbH
    • Correspondent: Robert A. Schaef, Harman International Industries, Incorporated, 2000 Research Drive, Stamford, CT, 06901, US. This correspondent recurs in this chain.
    • Context: Inventor assigned patent rights to their employer.
  • 2007-12-06 (executed) / recorded 2008-04-24 — Reel 020613/0287

    • Conveyance: Assignment
    • Assignor: Andreas Löw
    • Assignee: Harman Becker Automotive Systems GmbH
    • Correspondent: Robert A. Schaef, Harman International Industries, Incorporated, 2000 Research Drive, Stamford, CT, 06901, US. This correspondent recurs in this chain.
    • Context: Inventor assigned patent rights to their employer.
  • 2010-01-19 (executed) / recorded 2010-02-18 — Reel 024040/0173

    • Conveyance: Assignment
    • Assignor: Harman Becker Automotive Systems GmbH
    • Assignee: Nuance Communications, Inc.
    • Correspondent: Patrick F. O'Reilly, JR., c/o Nuance Communications, Inc., 1 Wayside Road, Burlington, MA, 01803, US.
    • Context: Transfer of patent rights as part of an asset purchase agreement.
  • 2019-10-23 (executed) / recorded 2019-10-23 — Reel 050836/0191

    • Conveyance: Assignment
    • Assignor: Nuance Communications, Inc.
    • Assignee: Cerence Inc.
    • Correspondent: David D. Lesh, Lempia Summerfield & Katz, LLC, One North LaSalle Street, Suite 2900, Chicago, IL, 60602, US.
    • Context: Transfer of intellectual property as part of an intellectual property agreement.
  • 2019-10-29 (executed) / recorded 2019-10-29 — Reel 050860/0677

    • Conveyance: Corrective Assignment
    • Assignor: Nuance Communications, Inc.
    • Assignee: Cerence Operating Company
    • Correspondent: David D. Lesh, Lempia Summerfield & Katz, LLC, One North LaSalle Street, Suite 2900, Chicago, IL, 60602, US. This correspondent recurs in this chain.
    • Context: Corrective assignment to update the assignee name.
  • 2019-11-07 (executed) / recorded 2019-11-07 — Reel 050901/0571

    • Conveyance: Security Agreement
    • Assignor: Cerence Operating Company
    • Assignee: Barclays Bank PLC
    • Correspondent: Jennifer S. Blum, Esq., Goodwin Procter LLP, 620 Eighth Avenue, New York, NY, 10018, US.
    • Context: Securitization of assets.
  • 2020-06-12 (executed) / recorded 2020-06-12 — Reel 051648/0225

    • Conveyance: Release by Secured Party
    • Assignor: Barclays Bank PLC
    • Assignee: Cerence Operating Company
    • Correspondent: Jennifer S. Blum, Goodwin Procter LLP, 620 Eighth Avenue, New York, NY, 10018, US. This correspondent recurs in this chain.
    • Context: Release of security interest.
  • 2020-06-15 (executed) / recorded 2020-06-15 — Reel 051670/0410

    • Conveyance: Security Agreement
    • Assignor: Cerence Operating Company
    • Assignee: Wells Fargo Bank, N.A.
    • Correspondent: Jennifer S. Blum, Goodwin Procter LLP, 620 Eighth Avenue, New York, NY, 10018, US. This correspondent recurs in this chain.
    • Context: Securitization of assets.
  • 2022-04-19 (executed) / recorded 2022-04-19 — Reel 053074/0231

    • Conveyance: Corrective Assignment
    • Assignor: Nuance Communications, Inc.
    • Assignee: Cerence Operating Company
    • Correspondent: David D. Lesh, Lempia Summerfield & Katz, LLC, One North LaSalle Street, Suite 2900, Chicago, IL, 60602, US. This correspondent recurs in this chain.
    • Context: Corrective assignment to replace a previous conveyance document.
  • 2025-01-02 (executed) / recorded 2025-01-02 — Reel 057037/0746

    • Conveyance: Release
    • Assignor: Wells Fargo Bank, National Association
    • Assignee: Cerence Operating Company
    • Correspondent: Kevin M. O'Brien, Womble Bond Dickinson (US) LLP, One Wells Fargo Center, 301 South College Street, Suite 3500, Charlotte, NC, 28202, US.
    • Context: Release of security interest.

Timeline diagram

timeline
    title Ownership of US 8306815
    2007 : Inventors assign to Harman Becker
    2010 : Assigned to Nuance Comm Inc
    2012 : Issued
    2019 : Assigned to Cerence Inc
         : Corrective to Cerence Operating Co
         : Security agreement to Barclays
    2020 : Release by Barclays
         : Security agreement to Wells Fargo
    2022 : Corrective from Nuance to Cerence
    2025 : Release by Wells Fargo

NPE / troll-pattern signals

  1. Shell-entity transfernot present. All identified assignees (Harman Becker Automotive Systems GmbH, Nuance Communications Inc., Cerence Inc., Cerence Operating Company) are operating companies in the automotive or speech technology sectors.
  2. Known asserter in the chainnot present. None of the assignees (Harman Becker Automotive Systems GmbH, Nuance Communications Inc., Cerence Inc., Cerence Operating Company) are listed as known NPEs or high-frequency plaintiffs in public databases like RPX or Unified Patents.
  3. Repeat correspondent across the chainpresent. Robert A. Schaef (Harman International Industries, Incorporated) appears on reels 020613/0285, 020613/0286, and 020613/0287. David D. Lesh (Lempia Summerfield & Katz, LLC) appears on reels 050836/0191, 050860/0677, and 053074/0231. Jennifer S. Blum (Goodwin Procter LLP) appears on reels 050901/0571, 051648/0225, and 051670/0410. The recurrence of these correspondents across multiple related assignments, especially with Cerence, is noted.
  4. Cascading transfersnot present. The transfers involve corporate acquisitions or internal re-organizations, not rapid transfers through multiple shell LLCs.
  5. Pre-litigation transferunclear. No litigation is publicly known for this patent, so this signal cannot be assessed.
  6. Bankruptcy fire-salenot present. There is no indication of any assignee entering bankruptcy proceedings leading to a patent sale.
  7. Privateeringnot present. There is no evidence from SEC filings or other public reports suggesting a privateering arrangement.
  8. Defensive aggregator (anti-NPE)not present. The chain does not terminate at a known defensive aggregator.

Verdict

Operating-company assertion (current assignee ships products embodying the claims and is suing actual competitors). The assignment history shows transfers between established operating companies (Harman Becker, Nuance, Cerence) in the speech technology and automotive sectors. Cerence Operating Company, the current assignee, is a direct successor to Nuance's automotive division and develops products embodying speech dialog systems.

Verification: https://assignmentcenter.uspto.gov/patent/index.html

Generated 5/29/2026, 9:04:10 PM

Prior art

Earlier patents, publications, and products that may anticipate or render the claims unpatentable.

✓ Generated

Based on a review of the citations for US patent 8,306,815, the following prior art references are identified as most relevant to the patent's claims. The analysis focuses on the potential for anticipation under 35 U.S.C. § 102, which requires a single prior art reference to disclose each and every element of a claimed invention.

The core inventive concept of US 8,306,815, particularly in its independent claims, involves a specific dual-feedback control architecture:

  1. An analysis signal, containing non-semantic information about the speech input (e.g., noise, pitch, location), is used to control a speech output unit.
  2. A recognition result, containing the semantic content (the recognized words), is used to control the signal pre-processor unit.

The following prior art references touch upon key elements of this system but do not appear to disclose the complete, specific architecture claimed.


1. US 2004/0064315 A1

  • Full Citation: US Patent Application Publication No. 2004/0064315 A1, "Acoustic confidence driven front-end preprocessing for speech recognition in adverse environments."
  • Publication Date: April 1, 2004
  • Brief Description: This document describes a speech recognition system that calculates an "acoustic confidence measure" based on the signal-to-noise ratio (SNR) of an input audio signal. This measure, which is a form of non-semantic analysis, is then used to adaptively control the front-end pre-processing of the signal. For example, more aggressive noise reduction can be applied when acoustic confidence is low.
  • Potential Anticipation of Claims:
    • This reference discloses using a non-semantic analysis of an input signal (the "acoustic confidence measure," analogous to the "analysis signal" in patent 8,306,815) to control a signal pre-processor.
    • This anticipates the general concept of adapting pre-processing based on signal quality. However, it appears to differ from the specific architecture of Claim 1, which requires the analysis signal to control the output unit and the recognition result to control the pre-processor.
    • It does not appear to disclose the complete feedback loop structure of independent claims 1, 20, 22, and 23.

2. US 5,524,169 A

  • Full Citation: US Patent No. 5,524,169 A, "Method and system for location-specific speech recognition," assigned to International Business Machines Corporation (IBM).
  • Issue Date: June 4, 1996
  • Brief Description: This patent details a system that improves speech recognition by using the speaker's physical location. The system determines where the user is (e.g., driver vs. passenger seat in a car) and selects a location-specific acoustic model or grammar to process the speech, thereby increasing accuracy.
  • Potential Anticipation of Claims:
    • This reference strongly anticipates the use of speaker location, a non-semantic characteristic, to control a speech dialog system. This is directly relevant to dependent claims 8 and 9 of patent 8,306,815.
    • It also anticipates the concept of using this analysis to control the speech recognition unit, as claimed in dependent claims 16 and 17.
    • However, US 5,524,169 does not appear to teach using the location information to control the speech output unit, nor does it disclose using the semantic recognition result to control the signal pre-processor. Therefore, it does not anticipate the complete system defined in independent claims 1, 20, 22, or 23.

3. US 7,016,837 B2

  • Full Citation: US Patent No. 7,016,837 B2, "Voice recognition system," assigned to Pioneer Corporation.
  • Issue Date: March 21, 2006
  • Brief Description: This patent describes a system that adapts to both the environment and the user's speaking style. It analyzes non-semantic vocal characteristics like speaking speed and volume. This analysis is used to select the most appropriate acoustic model from a library, thereby improving recognition when a user's speaking style changes.
  • Potential Anticipation of Claims:
    • This reference teaches the analysis of non-semantic characteristics like volume, as specified in dependent claim 10 of patent 8,306,815.
    • It uses this analysis to control the speech recognition unit by selecting models, which is the subject of dependent claims 16 and 17.
    • The patent does not appear to disclose the key control loops of the independent claims: using the non-semantic analysis to control the system's output and using the recognized words to control the system's input pre-processing. It therefore fails to anticipate independent claims 1, 20, 22, or 23.

4. US 2003/0061037 A1

  • Full Citation: US Patent Application Publication No. 2003/0061037 A1, "Method and apparatus for identifying noise environments from noisy signals."
  • Publication Date: March 27, 2003
  • Brief Description: This application discloses a system that identifies the specific type of background noise in a speech signal (e.g., car noise, babble). It then uses this noise identification to select an optimal model for speech recognition or noise compensation, with the goal of improving recognition accuracy.
  • Potential Anticipation of Claims:
    • This reference provides a clear example of generating an "analysis signal" based on a non-semantic characteristic (the noise environment), as described in dependent claim 4 of patent 8,306,815.
    • It further uses this analysis to adapt the recognition process, which is relevant to dependent claims 16 and 17.
    • As with the other references, it does not describe the specific dual-control architecture required by the independent claims. It lacks the teaching of using the noise analysis to control the speech output unit and using the recognized words to control the signal pre-processor. Thus, it does not anticipate claims 1, 20, 22, or 23.

Generated 5/8/2026, 10:04:22 PM

Obviousness

Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.

✓ Generated

Based on the provided prior art, an analysis of the obviousness of US patent 8,306,815 under 35 U.S.C. § 103 is as follows.


Obviousness Analysis of US Patent 8,306,815

A determination of obviousness under 35 U.S.C. § 103 requires assessing whether the differences between the claimed invention and the prior art are such that the invention as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art (PHOSITA). In this context, a PHOSITA would be an engineer or computer scientist with expertise in digital signal processing and speech recognition systems.

The core inventive concept of the '815 patent, as articulated in the independent claims (1, 20, 22, and 23), is a speech dialog system with a dual-feedback control architecture. This architecture involves two distinct control loops:

  1. Analysis-to-Output Loop: Using non-semantic information from a speech input (the analysis signal) to control the system's speech output.
  2. Recognition-to-Input Loop: Using the semantic meaning of the recognized words (the recognition result) to control the system's front-end signal pre-processing.

An argument for obviousness can be constructed by combining the teachings of the cited prior art references, as each reference discloses a key piece of the overall system. A PHOSITA would have been motivated to combine these teachings to solve the well-known problem of improving the robustness and usability of speech dialog systems in noisy and dynamic environments, such as inside a vehicle.


Combination Rendering Independent Claims 1, 22, and 23 Obvious

A combination of US 2004/0064315 A1 (Deisher) and US 5,524,169 A (IBM), supplemented with common knowledge in audio engineering and user interface design, would render the subject matter of claims 1, 22, and 23 obvious.

1. The Analysis-to-Output Loop:

  • What the Prior Art Teaches: Deisher teaches the generation of an "acoustic confidence measure" based on the signal-to-noise ratio (SNR) of an input signal. This is directly analogous to the analysis signal of the '815 patent containing information on the noise component (claim 4). Deisher uses this analysis to control the pre-processor.
  • Motivation to Combine/Modify: A primary challenge in noisy environments is ensuring the user can hear the system's feedback. If the system analyzes the background noise and determines it to be high (as taught by Deisher), a PHOSITA would be motivated to ensure the dialog is not broken by the user's inability to hear the system's response. The most direct and obvious solution to this problem is to increase the volume of the speech output unit. This is a basic principle of audio engineering and user interface design. Therefore, a PHOSITA would find it obvious to take the noise analysis taught by Deisher and use it to control the output volume, thereby creating the Analysis-to-Output Loop. The motivation is to improve the reliability of the human-machine interaction, a predictable outcome. This would render the feature of controlling the speech output unit based on the analysis signal (as required by claim 1) obvious.

2. The Recognition-to-Input Loop:

  • What the Prior Art Teaches: The prior art, including IBM and Pioneer, teaches using context to modify the speech recognition process. IBM specifically teaches using speaker location to select appropriate acoustic models. While this isn't a direct Recognition -> Input loop, it establishes the principle of using contextual information to adapt the system.
  • Motivation to Combine/Modify: A PHOSITA would understand that the semantic content of a recognized command can itself provide powerful context about future acoustic conditions. For example, if a user in a car says, "Open the driver's side window," the recognition result contains a semantic prediction that the noise environment is about to change dramatically. To maintain the system's robustness for the next command, a PHOSITA would be motivated to use this predictive information. Deisher teaches having an adaptive pre-processor. It would have been an obvious step to feed the semantic context from the recognized words into the control logic for Deisher's adaptive pre-processor. For instance, upon recognizing "open window," the system could proactively adjust the noise cancellation filter coefficients in anticipation of new wind noise. This creates the Recognition-to-Input Loop. The motivation is to improve future recognition performance in a dynamically changing environment, which is a predictable improvement, not an inventive leap.

Combining these two motivated modifications—using noise analysis to control output volume and using semantic recognition to control the input pre-processor—discloses all the key elements of independent claims 1, 22, and 23.


Combination Rendering Independent Claim 20 Obvious

Independent claim 20 is narrower, focusing only on the Recognition-to-Input Loop: "a speech dialog control unit configured to receive the recognition result, the speech dialog control unit configured to control the signal pre-processor unit based upon the recognition result."

This claim is rendered obvious by the combination of a standard speech recognition system with the adaptive pre-processor taught by Deisher (US 2004/0064315 A1).

  • Deisher teaches a signal pre-processor with adjustable parameters that are controlled based on an "acoustic confidence measure."
  • A PHOSITA would be motivated to find additional sources of information to improve the control of this pre-processor, especially in dynamic environments. As described above, the semantic content of a recognized phrase (e.g., "turn on the fan," "call mom on speakerphone") provides a clear prediction of an impending change to the acoustic environment (fan noise, echo path changes).
  • It would have been obvious to a PHOSITA to use the output of the speech recognizer as an additional input to the control logic for the pre-processor. The motivation is clear: to preemptively adapt the noise and echo cancellation filters before the acoustic environment changes, thus improving the recognition accuracy of the next utterance. This is a predictable optimization, not an invention. This combination directly teaches the control loop claimed in claim 20.

Conclusion:

While no single prior art reference discloses the specific dual-feedback architecture of US 8,306,815, the constituent elements existed in the art. A person of ordinary skill in the art, faced with the well-known problem of making speech dialog systems work better in cars and other noisy environments, would have been motivated to combine these known elements. Using noise analysis to control output volume and using recognized commands to predictively adjust input filtering are obvious engineering solutions that yield predictable improvements in system usability and robustness. Therefore, the independent claims of US patent 8,306,815 would be found obvious under 35 U.S.C. § 103.

Generated 5/8/2026, 10:07:33 PM

Extensions

Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.

✓ Generated

Detailed Analysis of U.S. Patent 8,306,815

Patent Term Adjustments (PTA) and Extensions (PTE):

A search of the USPTO's public records for U.S. Patent 8,306,815 indicates that there have been no Patent Term Adjustments (PTA) or Patent Term Extensions (PTE) granted for this patent. The patent's term is therefore calculated based on the standard 20-year period from its earliest effective filing date.

Continuity Data: Continuation and Divisional Applications:

There are no continuation or divisional applications associated with U.S. Patent 8,306,815. The patent application, serial number 11/951,904, stands as a single, non-provisional application.

Related Patent Family Members:

U.S. Patent 8,306,815 claims priority to European Patent Application No. 06025974.4, filed on December 14, 2006. This establishes a direct relationship within its patent family. Other international counterparts may exist but are not immediately evident from the U.S. prosecution history.

Projected Expiration Date:

The filing date of the U.S. application (11/951,904) is December 6, 2007. The standard term for a utility patent filed after June 8, 1995, is 20 years from the earliest U.S. non-provisional filing date. In this case, since the U.S. application claims priority to an earlier European application, the 20-year term is calculated from the U.S. filing date.

Therefore, the projected expiration date for U.S. Patent 8,306,815 is December 6, 2027. This calculation is based on the filing date and the absence of any term adjustments. It is important to note that the enforceability of a patent can also be affected by the timely payment of maintenance fees, which are due at 3.5, 7.5, and 11.5 years after the patent grant date. A review of the USPTO's maintenance fee records would be necessary to confirm the patent's current status.

Generated 5/8/2026, 10:04:18 PM

Derivative works

Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.

✓ Generated

Defensive Disclosure for US Patent 8,306,815

Publication Date: 2026-05-08
Subject: Derivatives and obvious-in-part improvements to U.S. Patent 8,306,815, "Speech dialog control based on signal pre-processing." This document is intended to enter the public domain as prior art.

This disclosure details a series of derivative inventions, applications, and component substitutions related to the core claims of US Patent 8,306,815 (hereafter 'the '815 patent'). The purpose is to preemptively render obvious any future patent claims on these incremental variations. The core inventive concept of the '815 patent is a speech dialog system where a signal pre-processor generates both an enhanced speech signal (for recognition) and an analysis signal (containing non-semantic characteristics). A control unit uses the analysis signal to manage outputs and the recognition result to manage the pre-processor itself.


1. Material & Component Substitution Derivatives

1.1. Neuromorphic Co-Processor for Analysis Signal Generation

Enabling Description: This variation replaces the general-purpose processor or DSP described for the signal pre-processor unit (202) with a dedicated neuromorphic processing unit (NPU) based on spiking neural network (SNN) architecture. The NPU is specifically tasked with generating the analysis signal. The microphone array input is converted into a series of asynchronous temporal spikes. The NPU's inherent parallelism and event-driven nature allow it to compute non-semantic characteristics like pitch, volume, and stationarity with ultra-low power consumption and latency. For example, pitch is determined by the inter-spike interval of neural columns tuned to specific frequencies, and signal stationarity is determined by the temporal stability of spike patterns across the network. The speech dialog control unit (206) receives this NPU-generated analysis signal and controls the speech output unit as described in the '815 patent. The primary enhanced signal for the speech recognition unit (204) can still be generated by a traditional DSP.

graph TD
    A[Microphone Array] -->|Analog Signal| B(ADC);
    B -->|Digital Audio Stream| C{Signal Router};
    C --> D[DSP for Enhanced Signal];
    C --> E[Neuromorphic Processing Unit - NPU];
    D --> F[Speech Recognition Unit];
    E -->|Analysis Signal<br/>(Pitch, Volume, Stationarity)| G[Speech Dialog Control Unit];
    F -->|Recognition Result| G;
    G --> H[Speech Output Unit];
    G --> D;

1.2. Piezoelectric Polymer Film Sensors for Input

Enabling Description: The standard microphone input unit (104) is replaced with an array of piezoelectric polymer film sensors, such as polyvinylidene fluoride (PVDF), embedded directly into the surfaces of a vehicle's interior cabin. These sensors detect speech as mechanical vibrations propagating through the solid structures. The signal pre-processor unit (202) is equipped with a specific transfer function model to translate the structural vibration data into an intelligible acoustic signal. The key advantage is that the analysis signal can now include non-semantic information about physical interactions, such as a passenger tapping the dashboard or the vibration signature of a window being open. The speech dialog control unit (206) can use this richer analysis signal to, for example, differentiate between speech-directed noise and ambient environmental noise, and instruct the user to "Please stop tapping the dashboard" if it interferes with recognition.

sequenceDiagram
    participant User
    participant PVDF_Sensors
    participant PreProcessor
    participant ControlUnit
    User->>PVDF_Sensors: Speaks and Taps Dashboard
    PVDF_Sensors->>PreProcessor: Transmits Vibration Data
    PreProcessor->>PreProcessor: Generates Enhanced Speech Signal
    PreProcessor->>ControlUnit: Generates Analysis Signal (Speech + Tap Vibration)
    ControlUnit->>ControlUnit: Analyzes high non-stationarity from Tap
    ControlUnit->>User: Output: "Please stop tapping the dashboard"

1.3. Bone Conduction Transducer for Speech Output

Enabling Description: The output unit (106), typically a loudspeaker, is replaced by a bone conduction transducer integrated into the driver's headrest or seatbelt. This provides a private, non-interfering audio output that is inaudible to other passengers and does not create acoustic echo that must be cancelled by the signal pre-processor unit (202). The speech dialog control unit (206) receives the analysis signal indicating high ambient noise (as per claim 5). Instead of merely increasing the volume of a loudspeaker, which would further degrade the signal-to-noise ratio for the microphone, the control unit increases the vibrational amplitude of the bone conduction transducer, ensuring clear delivery of the synthesized speech output to the intended user without contributing to the acoustic noise floor.

flowchart TD
    subgraph System
        A(Input Signal) --> B{Pre-Processor};
        B -- Enhanced Signal --> C(Speech Recognition);
        B -- Analysis Signal <br/> (High Noise Detected) --> D{Control Unit};
        C -- Recognition Result --> D;
        D -- Control Signal --> E(Bone Conduction Transducer);
    end
    subgraph Environment
        F(High Ambient Noise) --> A;
    end
    E -- Vibrations --> G(User);

2. Operational Parameter Expansion Derivatives

2.1. Cryogenic Environment Application for Superconducting Electronics Monitoring

Enabling Description: The system is adapted to operate in a cryogenic environment (e.g., < 77 Kelvin) for monitoring superconducting quantum computing equipment. The input unit (104) is a specialized cryogenic acoustic sensor designed to detect minute acoustic emissions, such as flux "avalanches" or mechanical stresses, which are precursors to decoherence events. The signal pre-processor unit (202) runs algorithms specifically designed to filter the dominant noise source in this environment: the acoustic signature of the cryocooler system. The analysis signal quantifies non-semantic characteristics like the frequency and stationarity of high-frequency acoustic bursts indicative of a potential quench. The speech recognition unit (204) is repurposed to classify these acoustic signatures against a library of known failure modes. The speech dialog control unit (206) uses the analysis and classification to trigger an alert or initiate an automated system shutdown.

stateDiagram-v2
    [*] --> Monitoring
    Monitoring --> Quench_Detected: High non-stationary signal
    Quench_Detected --> Shutdown: Control unit initiates safe shutdown
    Shutdown --> [*]
    Monitoring --> Anomaly_Detected: Unclassified acoustic signature
    Anomaly_Detected --> Alert: Control unit alerts human operator
    Alert --> Monitoring

2.2. Hypersonic Flow Analysis in a Wind Tunnel

Enabling Description: The system is applied to analyze airflow in a hypersonic wind tunnel. The input unit is an array of high-frequency pressure transducers rated for >100 kHz operation. The system does not process "speech" but rather the "acoustic" pressure fluctuations associated with turbulent boundary layers and shockwave oscillations. The signal pre-processor unit (202) filters mechanical vibration from the tunnel infrastructure. The analysis signal is a multi-dimensional vector representing the non-semantic characteristics of the flow: stationarity (laminar vs. turbulent flow), dominant frequencies (flow instabilities), and spatial location of acoustic sources derived from beamforming the transducer array. The speech recognition unit is replaced by a pattern classifier that identifies specific aerodynamic events (e.g., "shockwave boundary layer interaction"). The control unit (206) uses the analysis signal and event classification to adjust wind tunnel parameters (e.g., angle of attack of the model) in real-time to maintain stable test conditions.

graph TD
    A[Pressure Transducer Array] --> B{Pre-Processor};
    B -- Filtered Pressure Data --> C[Aerodynamic Event Classifier];
    B -- Analysis Signal <br/>(Flow stationarity, source location) --> D{Tunnel Control Unit};
    C -- Event Classification <br/>(e.g., "Boundary Layer Separation") --> D;
    D --> E[Wind Tunnel Actuators <br/>(e.g., change model pitch)];

2.3. Deep-Sea Wellhead Integrity Monitoring

Enabling Description: The system is enclosed in a high-pressure housing and deployed on a subsea oil and gas wellhead (e.g., at 3000m depth, >30 MPa). The input unit is an array of hydrophones. The signal pre-processor is tuned to filter noise from ocean currents and distant marine life. The analysis signal quantifies non-semantic acoustic characteristics indicative of material fatigue or leaks, such as the high-frequency signature of gas escaping a fissure or the low-frequency creak of metal under stress. The speech recognition unit is replaced by a fault diagnostics classifier. The control unit uses the analysis signal (e.g., detecting a non-stationary, high-frequency hiss) and the classifier output ("Class-A Leak Detected") to actuate a safety valve on the wellhead and transmit an emergency alert to a surface control station via an acoustic modem.

sequenceDiagram
    participant Wellhead
    participant HydrophoneArray
    participant SubseaControlUnit
    participant SurfaceStation

    Wellhead->>HydrophoneArray: Begins leaking (high-frequency hiss)
    HydrophoneArray->>SubseaControlUnit: Provides raw acoustic data
    SubseaControlUnit->>SubseaControlUnit: Pre-processes signal, generates analysis signal (non-stationary hiss)
    SubseaControlUnit->>SubseaControlUnit: Classifies fault as "Leak"
    SubseaControlUnit->>Wellhead: Command: Close Safety Valve
    SubseaControlUnit->>SurfaceStation: Transmit Alert via Acoustic Modem

3. Cross-Domain Application Derivatives

3.1. Aerospace: Pilot Biometric Stress Monitoring

Enabling Description: The system is integrated into a fighter jet cockpit's communication system. The input unit is the pilot's helmet microphone. The signal pre-processor unit (202) is hardened against extreme vibration and electromagnetic interference. It generates an analysis signal from the pilot's speech that quantifies non-semantic biomarkers of stress and cognitive load, such as increased pitch (fundamental frequency), reduced pitch variability, and vocal tremor (low-frequency amplitude modulation). When the analysis signal indicates the pilot is under extreme duress (e.g., during a high-G maneuver), the speech dialog control unit (206) automatically simplifies the user interface, reading out only critical flight information via the speech output unit and temporarily disabling non-essential alerts to reduce cognitive load. The recognition result (e.g., "Eject, Eject, Eject") can be used to control the pre-processor to bypass all filtering, ensuring the command is recognized with minimal latency.

graph TD
    A[Pilot Speech] --> B{Pre-Processor};
    B -- Enhanced Signal --> C[Speech Recognition];
    B -- Analysis Signal <br/> (Pitch, Jitter, Shimmer) --> D{Dialog & UI Control Unit};
    D --> E[Cockpit Display System];
    D --> F[Speech Output Unit];
    C -- Recognition Result --> D;

    subgraph Biometric Feedback Loop
        D--Analysis shows high stress-->E(Simplify Display);
        D--Analysis shows high stress-->F(Prioritize Critical Audio Alerts);
    end

3.2. AgTech: Automated Livestock Health Screening

Enabling Description: In a smart barn, microphone arrays (input unit 104) continuously monitor the vocalizations of livestock (e.g., pigs or cattle). The signal pre-processor unit (202) filters out machinery and environmental noise. The analysis signal is generated to quantify non-semantic characteristics of the animal calls, such as volume, pitch, and stationarity, which are known to correlate with health states (e.g., a dry, low-amplitude cough is an early indicator of respiratory disease). The speech recognition unit (204) is trained not on words, but on a taxonomy of animal call types (e.g., cough, grunt, squeal). When the control unit (206) receives an analysis signal indicating a pathological cough and a recognition result confirming the classification, it can automatically trigger an external device (external device 112), such as a camera to focus on the specific animal and a gate to divert it into a pen for veterinary inspection.

flowchart LR
    A[Animal Vocalization] --> B(Microphone Array);
    B --> C{Pre-Processor};
    C --> D[Call Type Classifier<br/>(Cough, Squeal)];
    C --> E{Control & Sorting Unit};
    D -- Call Type: 'Cough' --> E;
    C -- Analysis Signal<br/>(Low Amplitude, High Stationarity) --> E;
    E -- Pathological Signature Detected --> F(Actuate Sorting Gate);
    E --> G(Log Event with Animal ID);

3.3. Finance: Algorithmic Trading based on Vocal Analysis

Enabling Description: The system is deployed to analyze the audio feed of quarterly earnings calls from publicly traded companies. The input unit (104) is a direct audio line from the webcast. The signal pre-processor unit (202) generates an analysis signal that quantifies the non-semantic vocal characteristics of the CEO and CFO, such as pitch variation, speech rate, and vocal fry, which are used as proxies for confidence and stress levels. This analysis signal is fed directly into an algorithmic trading model as a feature vector. The speech recognition unit (204) transcribes the call, and the speech dialog control unit (206) uses the recognition result to correlate the vocal stress markers with specific semantic content (e.g., vocal stress increases when the CEO says "we are confident in future guidance"). The trading algorithm uses this combined semantic and non-semantic data to execute trades.

sequenceDiagram
    participant Webcast
    participant System as '815 Derivative System'
    participant TradingAlgo

    Webcast->>System: Live audio of CEO speaking
    System->>System: Generate analysis signal (vocal stress markers)
    System->>System: Generate recognition result (transcription)
    System->>TradingAlgo: Stream analysis signal vector
    System->>TradingAlgo: Stream transcription text
    TradingAlgo->>TradingAlgo: Correlate high stress with phrase "future guidance"
    TradingAlgo->>TradingAlgo: Execute short-sell order

4. Integration with Emerging Tech Derivatives

4.1. AI-Driven Reinforcement Learning for Filter Optimization

Enabling Description: The speech dialog control unit (206) is augmented with a Reinforcement Learning (RL) agent. The RL agent's state is defined by the incoming analysis signal (noise level, pitch, etc.) and a confidence score from the speech recognition unit. Its action space consists of adjusting the parameters of the filters (e.g., noise suppression level, echo cancellation aggressiveness) within the signal pre-processor unit (202). The reward function is designed to maximize the speech recognition confidence score while minimizing processing latency. When a user confirms a command was correctly understood (e.g., by saying "yes" or through a physical action), the agent receives a positive reward, reinforcing the filter settings that led to success under those specific acoustic conditions. Over time, the system learns a sophisticated, context-aware policy for self-optimization.

graph TD
    A[Speech Input] --> B{Pre-Processor (State s_t)};
    B -- Enhanced Signal --> C{Speech Recognition};
    B -- Analysis Signal --> D{RL Agent (in Control Unit)};
    C -- Confidence Score --> D;
    D -- Action a_t (Adjust Filter Params) --> B;
    C -- Recognition Result --> E[Output/Action Execution];
    E -- External Feedback (Success/Failure) --> F(Reward Function r_t);
    F --> D;

4.2. IoT Sensor Fusion for Environmental Context

Enabling Description: The analysis signal is expanded to be a fused data stream, combining the acoustic non-semantic characteristics with data from external IoT sensors. In a smart home application, the speech dialog control unit (206) receives acoustic data from the microphone, but also receives data from a light sensor indicating the room is dark, a motion sensor indicating the user just entered, and a smart watch indicating the user's heart rate is elevated. The control unit uses this holistic context to interpret commands more accurately. For example, a low-volume utterance ("lights") combined with the IoT context of entering a dark room leads the system to immediately turn on the lights, whereas the same utterance in a bright, occupied room might prompt a clarifying question. The recognition result ("turn on the fan") would prompt the control unit to query a temperature sensor via the IoT network before adjusting the pre-processor, anticipating the acoustic noise the fan will introduce.

flowchart TD
    subgraph Acoustic
        A[Speech Input] --> B{Pre-Processor};
        B --> C[Speech Recognition];
    end
    subgraph IoT
        D[Light Sensor] --> E{Fusion Engine};
        F[Motion Sensor] --> E;
        G[Smart Watch] --> E;
    end
    B -- Analysis Signal --> E;
    C -- Recognition Result --> H{Control Unit};
    E -- Fused Context Signal --> H;
    H --> I[Control External Device];

4.3. Blockchain for Verifiable Command Auditing

Enabling Description: This derivative is for high-stakes environments like industrial process control or medical command systems. Every time a command is recognized, a transaction is created and committed to a private blockchain. The transaction payload contains the recognition result (e.g., "Increase reactor temperature to 500 Kelvin"), the full analysis signal (background noise level, speaker location, pitch), a cryptographic signature of the authenticated speaker, and a timestamp. The immutability of the blockchain provides a tamper-proof audit trail. The speech dialog control unit (206) is also the blockchain client. Before executing a critical command from an external device (112), it first verifies that the transaction has been successfully committed to the ledger, ensuring there is a permanent record of the command and the precise context in which it was given.

sequenceDiagram
    participant User
    participant System
    participant Blockchain
    participant Critical_Device

    User->>System: Speaks command "Increase pressure"
    System->>System: Generates Recognition Result & Analysis Signal
    System->>Blockchain: Create Transaction (Result, Analysis, UserID, Timestamp)
    Blockchain-->>System: Confirmation (Tx Hash)
    System->>Critical_Device: Execute Command: Increase Pressure

5. "Inverse" or Failure Mode Derivatives

5.1. Failsafe Acoustic Watchdog Mode

Enabling Description: In this mode, the system is designed for safe failure. Upon detection of a processor fault or a software crash in the main speech dialog control unit (206), control is transferred to a simple, independent microcontroller. This microcontroller has access only to a single value from the signal pre-processor: the total signal energy (a minimalist analysis signal). The speech recognition unit and complex filtering are completely shut down. The microcontroller's sole function is to act as an acoustic watchdog: if the signal energy exceeds a predefined safety threshold for more than a set duration (indicating a potential feedback loop or system malfunction generating loud noise), it physically disconnects the output unit (106) via a relay switch to ensure the system fails silently and safely.

stateDiagram-v2
    state "Normal Operation" as Normal
    state "Failsafe Watchdog" as Failsafe
    state "Output Disconnected" as Disconnected

    [*] --> Normal
    Normal --> Failsafe : Processor Fault Detected
    Failsafe --> Disconnected : Signal Energy > Threshold for T > 2s
    Disconnected --> [*] : Manual Reset
    Failsafe --> [*] : Manual Reset

5.2. Privacy-Preserving Non-Semantic Control Mode

Enabling Description: This derivative is designed explicitly to not perform speech recognition to protect user privacy. The speech recognition unit (204) is either physically absent or disabled in firmware. The signal pre-processor unit (202) performs its function of generating an analysis signal based on non-semantic characteristics like volume, pitch, and the number of distinct speakers (determined via location beamforming). The speech dialog control unit (206) uses only this analysis signal to control ambient systems. For example, if the analysis signal indicates a rising conversation volume and multiple speakers, the control unit might automatically lower the volume of background music. If it detects a single, soft voice (low volume, single location), it might dim the lights for a more relaxed atmosphere. The system provides environmental control based on the texture of the conversation, without ever processing the content.

graph TD
    A[Acoustic Environment] --> B{Pre-Processor};
    B --x D(Speech Recognition - DISABLED);
    B -- Analysis Signal <br/> (Volume, Pitch, # of Speakers) --> C{Ambient Control Unit};
    C --> E[Control Lights];
    C --> F[Control Music Volume];

6. Combination Prior Art with Open-Source Standards

6.1. Combination with WebRTC Standard

  • Scenario: A browser-based customer service application using WebRTC for real-time audio chat.
  • Implementation: The '815 patent's signal pre-processor unit and control unit are implemented in JavaScript using the Web Audio API. The analysis signal is augmented with non-semantic data from the WebRTC RTCPeerConnection.getStats() method, which provides metrics like jitter, packetsLost, and roundTripTime. When the analysis signal indicates high network jitter and high local background noise, the control unit not only increases the volume of the agent's voice but also sends a command via the WebRTC data channel to the agent's client, suggesting they speak more slowly to improve intelligibility over the poor connection.

6.2. Combination with MQTT Protocol

  • Scenario: A voice-controlled factory automation system using the lightweight MQTT publish/subscribe protocol.
  • Implementation: A worker wearing a headset (input unit) issues commands. The signal pre--processor unit on the worker's wearable device publishes the analysis signal to an MQTT topic (e.g., factory/zone3/worker1/audio/analysis) and the enhanced speech signal to another topic. A server-based speech recognition unit (using an open-source engine like Vosk) subscribes to the enhanced signal topic and publishes the recognition result to a results topic. The control unit subscribes to the analysis and results topics. If it receives a command ("activate press") and the analysis signal indicates a very high noise level (suggesting the worker is right next to a loud machine), it will publish a command to a feedback output unit on the worker's headset: "Confirmation required. Are you at a safe distance from the press?" before relaying the command to the machine's MQTT topic.

6.3. Combination with Kaldi ASR Toolkit

  • Scenario: An in-vehicle infotainment system where the speech recognition unit is an embedded instance of the open-source Kaldi toolkit.
  • Implementation: The speech dialog control unit leverages the rich, detailed output from Kaldi, which is more than just the final text. The recognition result includes the full recognition lattice (a graph of alternative word hypotheses) and per-word confidence scores. The control unit implements the feedback loop described in claim 20. When a user says "I'm going to open the sun roof," the recognition result is fed to the control unit. The control unit, based on the semantic meaning of "sun roof," sends a command to the signal pre-processor unit to proactively adjust its adaptive noise cancellation filter coefficients, anticipating the specific change in wind noise profile that opening the sunroof will create. This predictive control, triggered by the Kaldi-recognized phrase, improves the robustness of subsequent speech recognition.

Generated 5/8/2026, 10:05:16 PM

Keep exploring

Other patents in High-Tech (T)

See all High-Tech (T) patents →