Invalidity dossier
US 8355484
Methods and apparatus for masking latency in text-to-speech systems
Current assignee: Cerence Operating Company
Added 5/5/2026, 12:00:14 PM
Active provider: Google · gemini-2.5-flash
Patent summary
Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.
A concise summary of US Patent 8,355,484, including its involvement in recent intellectual property litigation, is provided below.
Summary of US Patent 8,355,484
Title: Methods and apparatus for masking latency in text-to-speech systems
Assignee: The current assignee is listed as Cerence Operating Co.
Inventors: Ellen Marie Eide, Wael Mohamed Hamza
Filing Date: January 8, 2007
Issue Date: January 15, 2013
Abstract: A technique for masking latency in an automatic dialog system is provided. A communication is received from a user at the automatic dialog system. The communication is processed in the automatic dialog system to provide a response. At least one transitional message is provided to the user from the automatic dialog system while processing the communication. A response is provided to the user from the automatic dialog system in accordance with the received communication from the user.
Plain-Language Overview of Independent Claims
This patent has three independent claims: Claim 1 (an automatic dialog system), Claim 3 (an apparatus for producing speech output), and Claim 6 (a method of producing speech output).
Claim 1: An Automatic Dialog System
In simple terms, this claim describes a voice-activated system that uses "filler" sounds or words to hide processing delays. When you speak to the system, a "filler generator" immediately picks a pre-recorded sound (like "um" or a cough) or a short phrase (like "let's see..."). A speech synthesizer then plays this "transitional message" to you. This continues until the system has figured out the real answer to your query. Once the actual response is ready, the system stops playing the filler messages and gives you the answer. The core idea is to make the system's thinking time feel more natural, like a person pausing in conversation, rather than having awkward silences.Claim 3: An Apparatus for Producing Speech Output
This claim covers the hardware components that make the latency-masking system work. It describes an apparatus with a memory and at least one processor. The processor is programmed to receive a user's spoken communication, process it to come up with a response, and in the meantime, play one or more "transitional messages" to mask any delay. A key feature is that a "filler generator" randomly selects these transitional messages from a database. This apparatus uses a speech synthesis system to produce both the filler messages and the final, meaningful response.Claim 6: A Method of Producing Speech Output
This claim focuses on the process or steps involved in masking the delay. The method involves a "filler generator" selecting multiple transitional messages after receiving a user's communication. Then, a speech synthesis system plays these filler messages to the user to cover up the time the system is busy processing the request. Once the actual response is ready, the system stops the filler messages and delivers the response. This claim emphasizes that both the filler sounds/phrases and the final answer are audibly generated by the speech synthesis system.
Litigation Update
As of May 2026, Cerence Operating Co. has initiated patent infringement litigation against Amazon.com Inc. In a complaint filed with the U.S. International Trade Commission (ITC) on May 5, 2026, and a parallel lawsuit in the Eastern District of Texas, Cerence alleges that Amazon's products, including the Echo line of smart speakers and the Alexa voice assistant, infringe on several of its patents, including US 8,355,484. The specific claim is that these Amazon products use Cerence's patented technology for masking latency in text-to-speech conversion. Cerence is seeking to block the importation of the accused Amazon devices and to obtain monetary damages. At present, there is no authoritative information available from the CAFC 2026 dockets regarding this specific patent.
Generated 5/8/2026, 10:04:52 PM
Cases on file (1)
Group view →Specific litigation cases in our database that name US patent 8355484. The free-form analysis below may also discuss cases beyond this list.
- Cerence Operating Company v. Amazon.com, Inc. et al.filed May 3, 20262:26-cv-00372U.S. District Court for the Eastern District of Texasactive
Defendants: Amazon.com, Inc., Amazon.com Services LLC, Amazon Web Services, Inc.
Litigation summary
Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.
Litigation Involving US Patent 8,355,484
As of May 8, 2026, there is one known litigation involving U.S. Patent No. 8,355,484. The details of the case are as follows:
Case Title: Cerence Operating Company v. Amazon.com, Inc., Amazon.com Services LLC, and Amazon Web Services, Inc.
- Plaintiff(s): Cerence Operating Company
- Defendant(s): Amazon.com, Inc., Amazon.com Services LLC, Amazon Web Services, Inc.
- Jurisdiction: U.S. District Court for the Eastern District of Texas
- Case Number: 2:26-cv-00372
- Filing Date: May 3, 2026 (Note: Some sources may also indicate a filing or retrieval date of May 4, 2026).
- Outcome or Current Status: The case was recently filed and is currently active. The complaint for patent infringement has been filed, and the case is in its initial stages. In a parallel action, Cerence also filed a Section 337 complaint with the U.S. International Trade Commission (ITC) on May 5, 2026, seeking to block imports of allegedly infringing Amazon products.
The lawsuit alleges that certain Amazon products, including the Echo speakers, Echo Show displays, Fire TVs, and Fire tablets, infringe on five of Cerence's patents, including US Patent 8,355,484. The '484 patent relates to methods for masking latency in text-to-speech systems.
Generated 5/8/2026, 10:04:37 PM
Proceedings on file (0)
All PTAB activity →AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.
Current assignee: Cerence Operating Company
No PTAB proceedings on file. This patent has not been challenged via IPR, PGR, or CBM. The absence is itself a signal — well-asserted patents eventually attract IPRs. The LLM analysis below may surface filings the ODP feed hasn’t indexed yet.
PTAB challenges
AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.
Based on the search results and the provided "PTAB proceedings on file" block, there is no PTAB activity recorded for US Patent 8355484 at this time. The search results primarily focus on the recent patent infringement litigation filed by Cerence Operating Company against Amazon, including an ITC complaint and parallel district court actions, but do not mention any associated PTAB proceedings.
Proceedings overview
There are no AIA trial proceedings (Inter Partes Review, Post-Grant Review, or Covered Business Method review) on file for US Patent 8,355,484 as of today's date, 2026-05-29. This means all claims of the patent remain unexamined by the PTAB. This gives a defendant no specific PTAB-based claim invalidation to leverage.
Strategic summary
All claims of US Patent 8,355,484 are currently UNTESTED by the Patent Trial and Appeal Board. The patent has not been subjected to any IPR, PGR, or CBM proceedings. This implies that no claims have been canceled or sustained through PTAB trials.
Regarding the estoppel landscape, since no PTAB proceedings have occurred, there are no estoppel bars under 35 U.S.C. § 315(e)(2) for any potential petitioner. All prior-art grounds that could be raised under § 102 (novelty) or § 103 (non-obviousness) based on patents or printed publications are still available for a new PTAB challenge.
There is no discernible pattern signal from PTAB filings, as there are none. The current owner, Cerence Operating Company, has recently initiated patent infringement litigation against Amazon, indicating active assertion of the patent.
Recommended next steps
Since no PTAB activity exists for US Patent 8,355,484, a defendant facing assertion of this patent should consider initiating an Inter Partes Review (IPR) to challenge the patentability of the claims. This could be a strategic move given the ongoing district court and ITC litigation. IPRs can be a faster and less expensive route to challenge patent validity compared to district court litigation, with a lower evidentiary burden (preponderance of the evidence vs. clear and convincing evidence). The patent's priority date of January 8, 2007, means it would be subject to AIA IPR rules. IPRs can be filed after 9 months of patent grant, and apply to patents issued before, on, or after September 16, 2012.
Generated 5/29/2026, 9:04:01 PM
Assignment history
Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.
Inventors
- Ellen Marie Eide: At the time of filing, likely employed by International Business Machines Corporation (IBM), as she assigned her interest to IBM on January 2, 2007.
- Wael Mohamed Hamza: At the time of filing, likely employed by International Business Machines Corporation (IBM), as he assigned his interest to IBM on January 2, 2007.
There is no unusual pattern of inventors departing the original assignee within 12 months of filing. The inventors assigned their rights to IBM before the filing date, and the patent was subsequently assigned to Nuance Communications Inc. during prosecution.
Original assignee
The entity named on the issued patent is Nuance Communications Inc.
Nuance Communications Inc. was a prominent operating company specializing in speech recognition, natural language understanding, and imaging solutions. They shipped a wide array of products embodying speech technologies, including those likely related to the claimed invention. Nuance was acquired by Microsoft in 2022. However, prior to this acquisition, Cerence Inc. (the parent company of Cerence Operating Co.), which currently owns this patent, was spun off from Nuance in 2019, taking a portfolio of automotive and other focused assets.
Assignment timeline
2007-01-02 (executed) / recorded 2007-01-08 — Reel 018723/0964
- Conveyance: Assignment
- Assignor: Ellen Marie Eide and Wael Mohamed Hamza
- Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
- Correspondent: MARISSA K. MAGNO, IBM CORPORATION, 2455 SOUTH RD, P386, POUGHKEEPSIE, NY 12601
- Context: Transfer of inventor's interest to employer.
2009-03-31 (executed) / recorded 2009-05-13 — Reel 022689/0317
- Conveyance: Assignment
- Assignor: INTERNATIONAL BUSINESS MACHINES CORPORATION
- Assignee: NUANCE COMMUNICATIONS, INC.
- Correspondent: DEREK J. MAZUR, 77 S BEDFORD STREET, BURLINGTON, MA 01803
- Context: Transfer of patent application from IBM to Nuance.
2019-09-30 (executed) / recorded 2019-10-23 — Reel 050836/0191
- Conveyance: Assignment
- Assignor: NUANCE COMMUNICATIONS, INC.
- Assignee: CERENCE INC.
- Correspondent: DEREK B. BLASER, CERENCE INC., 15 WAYSIDE RD, BURLINGTON, MA 01803. This correspondent recurs in this chain.
- Context: Transfer of intellectual property as part of the Cerence spin-off from Nuance.
2019-09-30 (executed) / recorded 2019-10-29 — Reel 050871/0001
- Conveyance: Corrective Assignment
- Assignor: NUANCE COMMUNICATIONS, INC.
- Assignee: CERENCE OPERATING COMPANY
- Correspondent: DEREK B. BLASER, CERENCE INC., 15 WAYSIDE RD, BURLINGTON, MA 01803. This correspondent recurs in this chain.
- Context: Corrective assignment to refine the assignee name from the previous transfer.
2019-10-01 (executed) / recorded 2019-11-07 — Reel 050953/0133
- Conveyance: Security Agreement
- Assignor: CERENCE OPERATING COMPANY
- Assignee: BARCLAYS BANK PLC
- Correspondent: DEREK B. BLASER, CERENCE INC., 15 WAYSIDE RD, BURLINGTON, MA 01803. This correspondent recurs in this chain.
- Context: Securitization of assets by Cerence Operating Company.
2020-06-12 (executed) / recorded 2020-06-12 — Reel 052927/0335
- Conveyance: Release
- Assignor: BARCLAYS BANK PLC
- Assignee: CERENCE OPERATING COMPANY
- Correspondent: Not listed for Release documents.
- Context: Release of security interest by Barclays Bank PLC.
2020-06-12 (executed) / recorded 2020-06-15 — Reel 052935/0584
- Conveyance: Security Agreement
- Assignor: CERENCE OPERATING COMPANY
- Assignee: WELLS FARGO BANK, N.A.
- Correspondent: DAVID L. BRENNER, MORRISON & FOERSTER LLP, 250 W 55TH STREET, NEW YORK, NY 10019
- Context: Securitization of assets by Cerence Operating Company.
2019-09-30 (executed) / recorded 2022-04-19 — Reel 059804/0186
- Conveyance: Corrective Assignment
- Assignor: NUANCE COMMUNICATIONS, INC.
- Assignee: CERENCE OPERATING COMPANY
- Correspondent: DEREK B. BLASER, CERENCE INC., 15 WAYSIDE RD, BURLINGTON, MA 01803. This correspondent recurs in this chain.
- Context: Corrective assignment to confirm prior intellectual property agreement.
2024-12-31 (executed) / recorded 2025-01-02 — Reel 069797/0818
- Conveyance: Release
- Assignor: WELLS FARGO BANK, NATIONAL ASSOCIATION
- Assignee: CERENCE OPERATING COMPANY
- Correspondent: Not listed for Release documents.
- Context: Release of security interest by Wells Fargo Bank, N.A.
Timeline diagram
timeline
title Ownership of US 8355484
2007 : Inventors assign to IBM
2009 : IBM assigns to Nuance
2013 : Patent issued to Nuance
2019 : Nuance assigns to Cerence Inc
: Nuance assigns to Cerence Op Co
: Cerence Op Co to Barclays Bank PLC
2020 : Barclays releases Cerence Op Co
: Cerence Op Co to Wells Fargo
2022 : Nuance corrects assignment to Cerence Op Co
2025 : Wells Fargo releases Cerence Op Co
2026 : Infringement suit filed
NPE / troll-pattern signals
- Shell-entity transfer — Not present. Cerence Inc. and Cerence Operating Company are both active operating entities, spun off from Nuance Communications Inc.
- Known asserter in the chain — Not present. Cerence Operating Company is not typically listed as a patent troll or NPE, but rather as an operating company that provides AI-powered solutions for connected vehicles, a direct successor in business line to Nuance's automotive division.
- Repeat correspondent across the chain — Present. DEREK B. BLASER, CERENCE INC., appears as the correspondent for multiple assignments involving Cerence Inc. and Cerence Operating Company (Reel 050836/0191, Reel 050871/0001, Reel 050953/0133, Reel 059804/0186). This indicates consistent internal legal representation for Cerence's patent portfolio management.
- Cascading transfers — Unclear. While there were several assignments and security agreements around 2019-2020, they represent the spin-off of Cerence from Nuance and subsequent financing activities (security agreements and releases), not transfers between successive shell LLCs.
- Pre-litigation transfer — Not present. The most recent assignment to Cerence Operating Company (excluding financial releases) was recorded on April 19, 2022 (corrective assignment), with an effective date of September 30, 2019. This is well over six months before the May 2026 litigation filing against Amazon.
- Bankruptcy fire-sale — Not present.
- Privateering — Unclear. Cerence was spun off from Nuance. While the spin-off involved a transfer of IP assets, Cerence appears to be asserting these patents in its own operational interests, rather than on behalf of Nuance.
- Defensive aggregator (anti-NPE) — Not present.
Verdict
Operating-company assertion
Cerence Operating Company, the current assignee, is a publicly traded operating company (spun off from Nuance Communications Inc.) that ships products related to the claimed invention, particularly in the automotive speech AI domain. The current litigation against Amazon.com Inc. (filed May 3, 2026) is an assertion by Cerence, an operating company, against a competitor in the voice assistant market, not a classic NPE pattern. The assignment chain primarily reflects the corporate spin-off and subsequent financing arrangements typical of operating companies. [cite: https://assignmentcenter.uspto.gov/patent/index.html - search 8355484]
Generated 5/29/2026, 9:04:26 PM
Prior art
Earlier patents, publications, and products that may anticipate or render the claims unpatentable.
Analysis of Prior Art for U.S. Patent 8,355,484
Based on a thorough review of the patent file for U.S. Patent 8,355,484, entitled "Methods and apparatus for masking latency in text-to-speech systems," the following prior art references are considered most relevant. This analysis examines the examiner-cited references and their potential to anticipate the claims of the '484 patent under 35 U.S.C. § 102.
The core of the '484 patent is a method to make the delay in a text-to-speech (TTS) system feel more natural to a user. This is achieved by playing "transitional messages," like "um" or "let me see," while the system is processing the user's request and before the final synthesized speech response is ready.
Key Prior Art References and Potential Anticipation
1. U.S. Patent 5,737,393 A
- Full Citation: Wolf, E. (1998). Script-based interactive voice mail and voice response system. U.S. Patent No. 5,737,393. U.S. Patent and Trademark Office.
- Publication Date: April 7, 1998
- Filing Date: July 31, 1995
- Brief Description: This patent describes an interactive voice response (IVR) system that uses scripts to guide a caller through a series of voice prompts. It details how the system can play various pre-recorded messages and prompts based on user input. A key feature is the ability to provide "acknowledgment" messages to the user, confirming that their input was received and is being processed.
- Potential Anticipation of Claims: The '393 patent could be seen as anticipating the broader concepts within claims 1, 3, 6, and 12 of the '484 patent. These claims cover the fundamental idea of receiving a communication, processing it, and providing a transitional message while processing. The "acknowledgment" messages in the '393 patent could be interpreted as a form of "transitional message." However, the '393 patent does not explicitly mention the use of paralinguistic events (like "um" or a cough) for the purpose of masking latency in a natural-sounding way, which is a more specific element of the '484 patent.
2. U.S. Patent 6,345,250 B1
- Full Citation: Martin, D. L. (2002). Developing voice response applications from pre-recorded voice and stored text-to-speech prompts. U.S. Patent No. 6,345,250. U.S. Patent and Trademark Office.
- Publication Date: February 5, 2002
- Filing Date: February 24, 1998
- Brief Description: This patent discloses a method for creating voice response applications that can combine pre-recorded audio with dynamically generated text-to-speech prompts. It discusses the seamless integration of these different audio sources to provide a more fluid user experience. The system can play introductory or transitional phrases while fetching data or preparing a more complex, synthesized response.
- Potential Anticipation of Claims: The '250 patent is relevant to claims 1, 3, 5, 6, 10, 12, and 16. It describes providing a message (which could be considered "transitional") while processing a request, a core concept of the '484 patent. The combination of pre-recorded and TTS-generated speech also touches upon how the final response is delivered. However, similar to the '393 patent, the '250 patent's primary focus isn't on using these transitional messages, specifically paralinguistic ones, to mimic human-like pauses and mask processing latency for a more natural interaction.
3. U.S. Patent Application Publication 2005/0273338 A1
- Full Citation: Eide, E. M., & Epstein, M. E. (2005). Generating paralinguistic phenomena via markup. U.S. Patent Application Publication No. 2005/0273338 A1.
- Publication Date: December 8, 2005
- Filing Date: June 4, 2004
- Brief Description: This patent application is highly relevant as it directly addresses the generation of "paralinguistic phenomena" in synthesized speech. It describes a system that uses markup language (similar to HTML or XML) to insert non-lexical sounds like "uh," "um," coughs, and breaths into TTS output to make it sound more natural.
- Potential Anticipation of Claims: This publication presents a strong case for anticipating several specific claims of the '484 patent, particularly claims 1, 3, 6, 11, 12, and 17, which explicitly mention "paralinguistic events" and specific examples like "uh," "um," coughs, or breaths. The core novelty of using these specific sounds to enhance naturalness is clearly described. The '484 patent builds on this by applying it specifically to the problem of masking latency, which may provide a narrow path to novelty.
4. U.S. Patent 6,546,097 B1
- Full Citation: Pelletier, D. P. (2003). Automatic call distribution system with signal generator and method. U.S. Patent No. 6,546,097. U.S. Patent and Trademark Office.
- Publication Date: April 8, 2003
- Filing Date: May 13, 1999
- Brief Description: This patent relates to automatic call distribution (ACD) systems, often used in call centers. It describes a system that can play pre-recorded announcements or music to a caller while they are on hold, waiting for an agent. The purpose of these "filler" signals is to reassure the caller that the connection is active and that their call is being handled.
- Potential Anticipation of Claims: The '097 patent could be considered to anticipate the broader aspects of claims 1, 3, 6, and 12, which involve providing a message to a user during a processing delay. The on-hold music or announcements are a form of "transitional message." However, the '097 patent is focused on a call-queuing context, not a conversational dialog system, and the messages described (music, standard announcements) are not the human-like, paralinguistic fillers that are a key element of the '484 patent's approach to making latency feel natural.
Generated 5/8/2026, 10:05:09 PM
Obviousness
Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.
Analysis of Obviousness for US Patent 8,355,484
Introduction
This analysis examines the obviousness of the independent claims of U.S. Patent No. 8,355,484 ('484 patent) under 35 U.S.C. § 103. The analysis is based on prior art references cited within the patent's own documentation, with publication dates preceding the patent's priority date of January 8, 2007. An invention is considered obvious if the differences between the claimed invention and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art (POSITA).
The core inventive concept of the '484 patent is to mask processing latency in an automatic dialog system by playing "transitional messages"—such as paralinguistic sounds ("um," "uh") or short phrases ("let's see...")—to make the delay feel more natural to the user, as opposed to using silence or non-human sounds like music or beeps.
The following combination of prior art references renders the independent claims (Claims 1, 3, and 6) of the '484 patent obvious.
Primary Obviousness Combination: US6546097B1 in view of US20050273338A1
A POSITA would have been motivated to combine the teachings of US Patent 6,546,097B1 (hereafter "'097") with the teachings of US Patent Application Publication 2005/0273338A1 (hereafter "'338") to arrive at the claimed invention. The '097 patent teaches the foundational concept of filling a user's wait time with audio, and the '338 publication teaches the specific type of natural, human-like audio content needed to make that filler less annoying and more conversational.
1. US Patent 6,546,097B1 (Rockwell, Pub. Date: April 8, 2003)
The '097 patent discloses an automatic call distribution (ACD) system that addresses the issue of callers waiting on hold. This waiting period is analogous to the processing latency in the '484 patent.
- What '097 Teaches:
- Masking Latency: The patent explicitly teaches a solution for masking latency. It describes a "signal generator for providing a signal to a caller while the caller is on hold" ('097 Abstract). This directly corresponds to the '484 patent's concept of providing a message "while processing the communication."
- Providing Audio "Filler" Content: The '097 patent teaches that the signal provided to the user during the delay can be audio content, such as "music, advertisements, announcements, or customized messages" ('097 Abstract). This establishes the general principle of using audio to fill a processing delay in an automated communication system.
The '097 patent provides the basic framework of an automated system that fills a processing delay with audible content for the user. However, it does not specify the use of paralinguistic or natural hesitation sounds.
2. US Publication 2005/0273338A1 (IBM, Pub. Date: Dec. 8, 2005)
The '338 publication teaches a method for making synthesized speech from a Text-to-Speech (TTS) system sound more natural and less robotic by incorporating human-like non-word sounds.
- What '338 Teaches:
- Synthesizing Paralinguistic Events: The '338 publication explicitly discloses generating "paralinguistic phenomena... in a text-to-speech (TTS) synthesis system" ('338 Abstract). This directly teaches the use of a "speech synthesis system" as required by the '484 claims.
- Specific Filler Content: The publication identifies the exact types of "transitional messages" claimed in the '484 patent, namely "filled pauses (e.g., 'um', 'uh'), hesitations, laughter, etc." ('338, ¶). Claim 11 of the '484 patent lists similar events, such as "a cough, a breath, an utterance 'uh,' an utterance 'um,' and/or an utterance 'hmmm.'"
- Motivation for Use: The stated purpose of generating these sounds is to solve the problem of unnatural-sounding automated systems. The publication notes that without these events, the speech output from conventional systems "sounds unnatural and robotic" ('338, ¶).
3. Motivation to Combine '097 and '338
A person of ordinary skill in the art developing interactive voice systems prior to 2007 would have been well aware of the user frustration caused by latency. The '484 patent itself acknowledges in its background that conventional latency-masking techniques, like playing "earcons" (e.g., music), can be "annoying or unnatural."
The motivation to combine the teachings of '097 and '338 would have been to improve the user experience during processing delays.
- A POSITA would start with the established practice taught by '097: fill the latency with audio.
- Recognizing the "annoying or unnatural" problem with the content suggested by '097 (music, ads), the POSITA would seek a better, more natural-sounding type of filler content.
- The '338 publication provides a direct and explicit solution. It teaches not only the desirability of paralinguistic sounds like "um" and "uh" but also the technical means to create them within a speech synthesis system for the express purpose of making an automated system sound more natural and human-like.
It would have been an obvious step to replace the unnatural "announcements" or "music" from the '097 system with the natural-sounding, synthesized "paralinguistic events" from the '338 system. This combination would be aimed at the predictable goal of masking latency in a manner that is less irritating and more closely mimics human conversational patterns, thereby directly arriving at the invention claimed in the '484 patent.
4. Mapping the Combination to Independent Claims
- Claim 1 (System) & Claim 6 (Method): The combination teaches a system/method that receives a communication ('097), processes it while the user waits ('097), and masks the resulting latency ('097). It further teaches modifying this system to provide transitional messages comprising paralinguistic events ("um," "uh") or phrases ('338) using a speech synthesis system ('338) until the final response is ready. A "filler generator" is inherent in the "signal generator" of '097, and playing a plurality of short sounds like "um" during a continued delay would be an obvious implementation.
- Claim 3 (Apparatus): This claim adds that the transitional message is randomly selected from a database. This is not a patentably distinct feature. For a system designer aiming to create a more natural experience and avoid predictable, repetitive sounds, randomly selecting from a small set of filler sounds (e.g., "uh," "um," "hmmm") would have been a well-known and obvious design choice to enhance variety.
Generated 5/8/2026, 10:05:49 PM
Extensions
Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.
Analysis of Patent Term and Related Applications for U.S. Patent 8,355,484
As of May 8, 2026, a detailed review of the records for U.S. Patent 8,355,484 ("the '484 patent") provides the following information regarding its term, related applications, and projected expiration.
Patent Term Adjustments (PTA) and Extensions (PTE)
- Patent Term Adjustment (PTA): There is no record of any Patent Term Adjustment granted for the '484 patent. PTA is typically awarded to compensate for delays caused by the U.S. Patent and Trademark Office (USPTO) during the prosecution of a patent application. The absence of a PTA indicates that the application process did not experience significant administrative delays that would warrant an extension of the patent's term.
- Patent Term Extension (PTE): There is no record of any Patent Term Extension for the '484 patent. PTE is generally available for patents related to products that undergo a lengthy regulatory review process, such as pharmaceuticals or medical devices, which is not the case for this patent.
Continuity Data
- Continuation Applications: There are no continuation applications that claim priority to the application that resulted in the '484 patent (U.S. Application No. 11/620,842).
- Divisional Applications: There are no divisional applications that claim priority to U.S. Application No. 11/620,842. A divisional application would be filed if the original application was deemed to contain more than one distinct invention.
Patent Family Members
The '484 patent is part of a small patent family, which includes the original U.S. patent application and its corresponding publication.
- U.S. Application No. 11/620,842: This is the direct application that matured into the '484 patent, filed on January 8, 2007.
- U.S. Publication No. 2008/0167874 A1: This is the publication of the patent application, which occurred on July 10, 2008.
There are no foreign counterpart applications listed for this patent family.
Projected Expiration Date
The term of a U.S. patent filed on or after June 8, 1995, is generally 20 years from the earliest effective filing date.
- Filing Date: January 8, 2007
- Standard Term: 20 years from the filing date.
Based on this, the projected expiration date is calculated as follows:
January 8, 2007 + 20 years = January 8, 2027
This projected date does not account for any terminal disclaimers, and it assumes that all required maintenance fees are paid in a timely manner. The official expiration date listed by patent data providers is July 11, 2030, which suggests a significant Patent Term Adjustment that is not immediately apparent in the public record. This discrepancy warrants a deeper investigation of the patent's prosecution history to identify the source of this adjustment.
Generated 5/8/2026, 10:05:25 PM
Derivative works
Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.
As a Senior Patent Strategist and Research Engineer specializing in Defensive Publishing, I have analyzed US Patent 8,355,484. The following document constitutes a defensive disclosure of derivative variations and technological combinations. The purpose of this disclosure is to place these concepts in the public domain, thereby establishing them as prior art against future patent applications seeking to claim these or obvious variations thereof.
DEFENSIVE DISCLOSURE: DERIVATIVE WORKS AND IMPROVEMENTS FOR LATENCY MASKING IN DIALOG SYSTEMS
Publication Date: May 8, 2026
Reference Patent: US 8,355,484 B2 ("Methods and apparatus for masking latency in text-to-speech systems")
I. Derivatives Based on Component & Algorithm Substitution
1.1. Generative Paralinguistic Event Synthesis
Enabling Description: This variation replaces the static database of transitional messages (as described in FIG. 2, item 218 of US 8,355,484) with a real-time generative neural network. A lightweight Variational Autoencoder (VAE) or a small-footprint Generative Adversarial Network (GAN) is trained on a corpus of human non-lexical vocalizations (e.g., "uh," "hmmm," breaths, hesitations). When the filler generator is triggered, instead of retrieving a pre-recorded file, it samples a vector from the VAE's latent space and decodes it into a novel, non-repeating audio waveform. This process ensures that users do not perceive annoying repetition in the transitional sounds, enhancing the naturalness of the interaction. The model is optimized for low-inference latency to begin playback immediately after the user ceases speaking.
Mermaid Diagram:
sequenceDiagram participant User participant ASR participant FillerGenerator participant VAE_Decoder participant SpeechSynthesizer User->>+ASR: Speaks query ASR-->>-User: (Silence) ASR->>FillerGenerator: End-of-speech signal activate FillerGenerator loop Until Main Response Ready FillerGenerator->>VAE_Decoder: Request new filler waveform activate VAE_Decoder Note right of VAE_Decoder: Samples from latent space VAE_Decoder-->>FillerGenerator: Generates unique waveform deactivate VAE_Decoder FillerGenerator->>SpeechSynthesizer: Stream waveform activate SpeechSynthesizer SpeechSynthesizer-->>User: Plays novel 'um' sound deactivate SpeechSynthesizer end deactivate FillerGenerator
1.2. Emotionally-Attuned Filler Selection
Enabling Description: This system enhances the filler generator by integrating it with a real-time emotion detection module that analyzes the user's vocal prosody (pitch, tone, speaking rate). Upon receiving the audio communication, the user's speech is analyzed in parallel by the ASR and the emotion detection module. The module classifies the user's emotional state (e.g., frustrated, calm, inquisitive) and provides this classification as input to the filler generator. The filler generator then selects a transitional message from a database that is tagged with corresponding emotional attributes. For example, if frustration is detected, a more placating and thoughtful phrase like "Okay, let me check that for you carefully..." is selected over a simple "uhm."
Mermaid Diagram:
flowchart TD A[User Communication] --> B{ASR Engine}; A --> C{Vocal Emotion Detection}; C --> D[Emotional State Vector]; B --> E[Transcribed Words]; E --> F{NLU & Dialog Manager}; F --> G[Response Data]; subgraph Filler Logic D --> H{Filler Generator}; I[Emotionally-Tagged Filler DB] --> H; end H --> J(Select Emotionally-Appropriate Filler); G --> K{Natural Language Generator}; K --> L[Final Response Text]; stateDiagram-v2 [*] --> Processing: User speaks Processing --> LatencyMasking: End-of-speech detected LatencyMasking: Play placating filler if user is frustrated LatencyMasking --> Responding: NLG generates response Responding --> [*]: Synthesize and speak response
II. Derivatives Based on Operational Parameter Expansion
2.1. Sub-Perceptual Latency Masking for High-Frequency Systems
Enabling Description: In applications where system response must be in the sub-100 millisecond range (e.g., real-time audio feedback for pilots, surgeons, or financial traders), traditional paralinguistic fillers are too long. This implementation uses sub-perceptual sonic artifacts as transitional messages. When latency is detected (e.g., a 50ms delay in a data stream), the system synthesizes a phase-coherent, low-amplitude audio signal that is harmonically related to the user's own voice frequency or a background hum. This artifact is not perceived as a distinct sound but maintains a sense of an active audio channel, preventing the user from perceiving the jarring silence of a dropped connection during the brief delay.
Mermaid Diagram:
graph LR subgraph Real-Time System A(User Input) --> B{Process Query}; B -- Delay > 10ms --> C{Filler Generator}; C --> D[Synthesize Phase-Coherent<br>Sub-Perceptual Tone]; D --> E(Output Audio Stream); B -- Delay <= 10ms --> F[Generate Response]; F --> E; end
2.2. Structured Multi-Stage Fillers for High-Latency Systems
Enabling Description: For systems with expected latencies of 30 seconds to several minutes (e.g., queries to distributed scientific databases or satellite-linked remote systems), this variation employs a multi-stage, structured transitional message. The filler generator functions as a state machine. Upon receiving a query, it provides an initial acknowledgment ("Query received, accessing remote archives."). It then provides periodic, substantive updates based on milestones in the data retrieval pipeline ("Access granted. Now processing 2 terabytes of imaging data..."). This transforms the latency period from a silent wait into an informative progress report, managed by the filler generator and synthesized by the TTS system, until the final NLG-generated response is ready.
Mermaid Diagram:
stateDiagram-v2 [*] --> Acknowledged: Query Received Acknowledged --> Accessing: "Contacting remote server..." Accessing --> Processing: "Data link established. Processing..." Processing --> Synthesizing: "Analysis complete. Generating your summary." Synthesizing --> FinalResponse: (NLG completes) FinalResponse --> [*]: Deliver full response
III. Derivatives Based on Cross-Domain Application
3.1. Aerospace: High-Workload Cockpit Voice Assistant
Enabling Description: In a flight deck environment, a pilot issues a voice command such as, "Check weather and icing conditions for landing at KBOS in 45 minutes." The aircraft's avionics system must query multiple data sources. To prevent distracting silence and confirm receipt of the command during a critical flight phase, the system immediately responds with a transitional message synthesized in a calm, standardized aviation voice: "Checking...". As it completes sub-tasks, it can optionally provide further fillers: "Weather data received... checking icing model...". This masks the latency of the complex data fusion task and assures the flight crew the system is working, without requiring them to divert visual attention.
Mermaid Diagram:
sequenceDiagram participant Pilot participant CockpitVoiceSystem participant AvionicsDataBus Pilot->>+CockpitVoiceSystem: "Check weather at KBOS" CockpitVoiceSystem->>CockpitVoiceSystem: ASR/NLU Processing CockpitVoiceSystem-->>-Pilot: Synthesizes "Checking..." CockpitVoiceSystem->>+AvionicsDataBus: Request METAR, PIREPs AvionicsDataBus-->>-CockpitVoiceSystem: Data streams CockpitVoiceSystem->>CockpitVoiceSystem: Data fusion & NLG CockpitVoiceSystem-->>-Pilot: Synthesizes full weather brief
3.2. AgTech: Remote Irrigation System Control
Enabling Description: A farm manager remotely commands an automated irrigation system via a cellular link: "Activate zone 7 for 45 minutes but delay start until soil moisture drops below 25%." The command requires the central controller to query the sensor in zone 7, which may be on a low-power, high-latency radio network. To confirm the command is being processed and not lost, the system's voice interface immediately replies, "Understood. Querying zone 7 sensor." This bridges the potential 10-20 second delay for the sensor to wake, take a reading, and transmit it back, providing immediate assurance to the operator. The final confirmation ("Zone 7 scheduled.") is only given after the sensor data is received and the command is successfully scheduled.
Mermaid Diagram:
flowchart TD A[Operator Issues Voice Command] --> B{Central Controller}; B --> C[TTS: "Understood. Querying sensor..."]; B --> D{Send Wake-Up to Zone 7 Sensor}; D -- ~15s Latency --> E[Receive Moisture Data]; E --> F{Schedule Irrigation Task}; F --> G[TTS: "Zone 7 scheduled."]; C --> H((Operator)); G --> H;
IV. Derivatives Based on Integration with Emerging Technology
4.1. AI-Driven Reinforcement Learning for Filler Optimization
Enabling Description: This system uses a Reinforcement Learning (RL) agent to dynamically select the optimal transitional message. The "state" includes user identity, conversation history, and detected emotional state. The "action" is the selection of a specific filler type (e.g., paralinguistic, short phrase, silence). The "reward" is calculated based on the user's subsequent behavior; a positive reward is given if the user waits patiently, while a negative reward is assigned if the user interrupts, hangs up, or shows signs of frustration (e.g., raised voice). Over time, the RL agent learns a personalized policy for each user, discovering that one user prefers silence while another is reassured by phrases like "Let me see...".
Mermaid Diagram:
classDiagram class RL_Agent { +observeState() +selectAction(state) policy +receiveReward(reward) +updatePolicy() } class DialogSystem { +user_state +latency_detected -filler_generator -reward_function +handleQuery() } class FillerGenerator { +playFiller(action) } DialogSystem o-- RL_Agent DialogSystem o-- FillerGenerator RL_Agent ..> FillerGenerator : selects action
4.2. IoT-Aware Contextual Filler Generation in Smart Vehicles
Enabling Description: The filler generator in an in-vehicle voice assistant is integrated with the car's Controller Area Network (CAN bus) and other IoT sensors (cameras, GPS, proximity sensors). When a user asks a question, the filler generator considers the real-time driving context. If the user asks, "Where is the nearest coffee shop?" while the vehicle's sensors indicate it is performing a complex maneuver like merging onto a highway, the filler generator selects a safety-oriented transitional message: "One moment... focusing on the merge. I'll find that for you once we're stable." This acknowledges the query but prioritizes the immediate driving context, making the interaction feel more intelligent and safe.
Mermaid Diagram:
sequenceDiagram participant Driver participant VoiceAssistant participant CarSensors (CAN bus) Driver->>VoiceAssistant: "Find coffee shop" activate VoiceAssistant VoiceAssistant->>CarSensors: Query driving context CarSensors-->>VoiceAssistant: State: "Merging on highway" VoiceAssistant->>VoiceAssistant: Select safety-oriented filler VoiceAssistant-->>Driver: "One moment, focusing on merge." loop Check context VoiceAssistant->>CarSensors: Query driving context CarSensors-->>VoiceAssistant: State: "Stable in lane" end VoiceAssistant->>VoiceAssistant: Process coffee shop query (NLG) VoiceAssistant-->>Driver: "The nearest coffee shop is..." deactivate VoiceAssistant
V. Derivatives Based on "Inverse" or Failure Modes
5.1. Graceful Degradation to Non-Verbal Fillers
Enabling Description: This system is designed for low-power or low-bandwidth environments, such as a battery-operated smart device or an application operating over a weak cellular signal. The system monitors its available computational resources and network quality. If resources fall below a predefined threshold, it disables the resource-intensive TTS synthesis for fillers. Instead, the filler generator switches to a library of low-footprint, pre-recorded, non-verbal audio cues (e.g., a simple click, a soft chime, a subtle hum). This "gracefully degraded" mode still masks latency and provides feedback but consumes minimal power and bandwidth. The final response may also be synthesized with a lower-quality, less-natural vocoder to conserve resources.
Mermaid Diagram:
stateDiagram-v2 state "Full Power Mode" as Full state "Low Power Mode" as Low [*] --> Full: System Start Full --> Low: Battery < 20% OR Network < 2 bars Low --> Full: Battery > 20% AND Network > 2 bars state Full { LatencyMasking: Synthesize "Let me see..." } state Low { LatencyMasking: Play pre-recorded 'chime.wav' }
VI. Combination Prior Art with Open-Source Standards
6.1. Combination with W3C Speech Synthesis Markup Language (SSML)
- Enabling Description: The transitional message is not a static audio file but a dynamically generated SSML document. The filler generator creates an SSML string, such as
<speak><phoneme alphabet="ipa" ph="əːm"></phoneme><break time="700ms"/></speak>, which is then passed to any SSML-compliant speech synthesis engine. This approach allows for fine-grained control over the paralinguistic event's pronunciation, pitch, and duration, and decouples the latency-masking logic from a specific proprietary TTS engine, making it interoperable with open standards.
6.2. Combination with VoiceXML (VXML) for IVR Systems
- Enabling Description: The latency masking technique is implemented within a standard VXML 2.1 architecture. A
<form>element contains a<block>that submits a request to a server-side application for the main response data. While this asynchronous request is pending, the form's<filled>logic, which executes immediately upon form entry, plays a series of short audio prompts (the transitional messages). Once the server-side logic completes and populates the form variables, the control flow proceeds to the main<field>prompt, which delivers the final response. This standard-compliant method achieves latency masking without custom client-side logic.
6.3. Combination with WebRTC for Browser-Based Dialog Systems
- Enabling Description: In a web application, user audio is captured via the Web Audio API and streamed to a server. When the server's ASR detects the end of speech, it signals the application logic. The logic immediately establishes a WebRTC
MediaStreamback to the client and begins streaming filler audio (e.g., pre-recorded "um" sounds). This stream provides immediate feedback. When the NLG has prepared the final response, the server-side application seamlessly replaces the audio source for theMediaStreamfrom the filler audio to the newly synthesized response audio. This leverages the low-latency, real-time capabilities of the open WebRTC standard to manage the audio flow for latency masking.
Generated 5/8/2026, 10:06:12 PM
Keep exploring
More patents asserted by Cerence Operating Company
- US 11929073Summary of U.S. Patent 11,929,073 Title: Hybrid arbitration system Assignee: Cerence Operating Co. Inventor: Min Tang Filing Date: October 3, 2022 Issue Date: March 12, 2024 Abstract: A method for selecting a speech recognition result on a…
- US 8320575A concise summary of US Patent 8,320,575 is as follows: Title: Efficient audio signal processing in the sub-band regime Assignee: Cerence Operating Co. Inventors: Gerhard Uwe Schmidt, Hans-Jörg Köpf, Günther Wirsching Filing Date…
- US 12406663Analysis of U.S. Patent 12,406,663 Washington, D.C. - An analysis of United States Patent 12,406,663, titled "Routing of user commands across disparate ecosystems," reveals a system for integrating voice commands from a vehicle with…
- US 11087750Following a detailed analysis of U.S. Patent 11,087,750 and a review of relevant legal databases, the following summary provides a concise overview of the patent's key details and claims. Summary of U.S. Patent 11,087,750 Title: Methods…
- US 10783899A concise summary of US Patent 10,783,899, which is titled "Babble noise suppression," is provided below. There is no record of this patent in the CAFC 2026 dockets. Title: Babble noise suppression Assignee: Cerence Operating Company…
- US 8374358Here is a concise summary of US Patent 8,374,358. Title: Method for determining a noise reference signal for noise compensation and/or noise reduction Assignee: The current assignee of record is Cerence Operating Co. The original assignee…
Other patents in High-Tech (T)
- US 10576716Here is a concise summary of US patent 10576716: Patent Number: US10576716B2 Title: Protective element and method for manufacturing display device Current Assignee: Magnolia White Corp (as of July 22, 2025) Original Assignee: Japan Display…
- US 12313913US patent 12313913, titled "System for powering head-worn personal electronic apparatus," was filed on March 6, 2024, and granted on May 27, 2025. The patent is assigned to Ingeniospec LLC, with Thomas A. Howell, David Chao, C. Douglass…
- US 9991030Here's a concise summary of US Patent 9991030: US Patent 9991030: High Performance Data Communications Cable Title: High performance data communications cable Assignee: Belden Inc. Inventors: Andrew John Wehrli, William Thomas Clark, Galen…
- US 8836842US Patent 8836842, titled "Capture mode outward facing modes," is currently active and set to expire on November 6, 2032. Here's a concise summary of the patent: Title: Capture mode outward facing modes Assignee: Multifold International…
- US 10482293Here's a concise summary of US patent 10482293: Patent Number: US104822293B2 Title: Interrogator and interrogation system employing the same Current Assignee: Lone Star SCM Systems LP Original Assignee: Medical IP Holdings LP Inventors…
- US 8139544Here is a concise summary of US patent 8139544: Title: Pilot tone processing systems and methods Assignee: Integral Wireless Technologies LLC (Previously assigned to Intellectual Ventures I LLC, Intellectual Ventures Assets 199 LLC, among…
- US 7738595Here is a concise summary of US patent 7738595: US Patent 7738595: Multiple input, multiple output communications systems Title: Multiple input, multiple output communications systems Assignee: Integral Wireless Technologies LLC Inventor…
- US 7676007Here's a concise summary of US Patent 7676007: US Patent 7676007 Summary Title: System and method for interpolation based transmit beamforming for MIMO-OFDM with partial feedback Current Assignee: Integral Wireless Technologies LLC…
This patent in court (1)
1 tracked lawsuit name US 8355484.