- Filed
- Mar 30, 2026
- Last modified
- Jul 21, 2026
- Petitioner
- Krisp Technologies, Inc.
- Inventor
- Ankita JHA et al
Invalidity dossier
US 12417756
Systems and methods for real-time accent mimicking
Current assignee: Krisp Technologies, Inc.
Added 4/30/2026, 3:11:01 PM
Active provider: Google · gemini-2.5-flash
Patent summary
Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.
Summary of U.S. Patent 12,417,756
As of the current date, April 30, 2026, the following is a summary of United States Patent number 12,417,756. This information is based on the authoritative patent text. A search of the United States Court of Appeals for the Federal Circuit (CAFC) 2026 dockets yielded no results for this patent number.
Title: Systems and methods for real-time accent mimicking
Assignee: Sanas Ai Inc.
Inventors:
- Ankita Jha
- Lukas Pfeifenberger
- Piotr Dura
- David Braude
- Alvaro Escudero
- Shawn Zhang
- Maxim Serebryakov
- Sharath Kashava Narayana
Filing Date: January 17, 2025
Issue Date: September 16, 2025
Abstract:
The patent describes technology for real-time accent mimicking. The system uses trained machine learning models to extract the accent features from a first user's speech. It then analyzes the speech of a second user to identify characteristics specific to their natural voice. A modified version of the second user's speech is then synthesized, which preserves the natural voice characteristics of the second user while mimicking the accent of the first user. This modified audio is then provided for output.
Plain-Language Overview of Independent Claims
U.S. Patent 12,417,756 has three independent claims, which define the core scope of the invention:
Claim 1: This claim protects a physical speech processing system. This system includes an audio interface, microphone, audio output, memory, and processors. The processors are instructed to:
- Use trained machine learning models to analyze a first person's speech to extract their specific accent features.
- Analyze a second person's speech to generate a profile of their unique "natural voice" characteristics (such as voice quality, phonetic patterns, and intonation).
- Create a new, modified version of the second person's speech that combines the first person's accent with the second person's preserved natural voice.
- Provide this modified speech as audio output.
Claim 7: This claim protects a method for real-time accent mimicking. This is a series of steps performed by a speech processing system, which includes:
- Applying machine learning models to extract accent features from a first user's speech.
- Analyzing a second user's speech to generate a "unique fingerprint" of their voice.
- Synthesizing a modified version of the second user's speech based on the extracted accent and the voice fingerprint.
- Providing the resulting audio data for output.
Claim 14: This claim protects a non-transitory computer-readable medium (such as a hard drive or other digital storage). This medium stores instructions that, when executed by a processor, cause the processor to perform the accent-mimicking process. The steps are functionally the same as the other independent claims: extracting a first user's accent features, generating characteristics of a second user's natural voice, synthesizing a modified version of the second user's speech that mimics the first accent while preserving the second user's voice, and providing the modified audio for output.
Legal Status Notes
The patent document indicates that a Post-Grant Review (PGR) proceeding has been initiated before the Patent Trial and Appeal Board (PTAB). The case number is cited as PGR2026-00033, and its status is listed as "Pending". This indicates that the validity of the patent's claims is currently under review by the USPTO.
Generated 4/30/2026, 4:30:17 PM
Cases on file (2)
Group view →Specific litigation cases in our database that name US patent 12417756. The free-form analysis below may also discuss cases beyond this list.
- Krisp Technologies, Inc. v. Sanas.AI Inc.filed Mar 30, 2026PGR2026-00033USPTO Patent Trial and Appeal Board (PTAB)Pending
Defendants: Sanas.AI Inc.
- Sanas.AI Inc. v. Krisp Technologies, Inc.filed Jul 7, 20253:25-cv-05666U.S. District Court for the Northern District of CaliforniaOngoing
Defendants: Krisp Technologies, Inc.
Litigation summary
Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.
Litigation and Administrative Challenges involving US Patent 12,417,756
As of April 30, 2026, U.S. Patent 12,417,756 is involved in one known district court litigation and one administrative challenge before the Patent Trial and Appeal Board (PTAB).
District Court Litigation
Case: Sanas.AI Inc. v. Krisp Technologies, Inc.
- Plaintiff: Sanas.AI Inc.
- Defendant: Krisp Technologies, Inc.
- Jurisdiction: U.S. District Court for the Northern District of California
- Case Number: 3:25-cv-05666
- Filing Date: July 7, 2025
- Status: The case is ongoing. On September 23, 2025, Sanas amended its original complaint to include infringement claims regarding U.S. Patent 12,417,756, which had been recently issued. The original complaint alleged theft of confidential information and infringement of other patents. According to a statement from Sanas, the court denied a motion from Krisp Technologies that attempted to invalidate Sanas's patents.
PTAB Administrative Challenge
In addition to the district court case, the patent is subject to a Post-Grant Review (PGR) proceeding, which challenges the patent's validity.
Case: PGR2026-00033
- Petitioner: Krisp Technologies, Inc.
- Patent Owner: Sanas.AI Inc.
- Jurisdiction: USPTO Patent Trial and Appeal Board (PTAB)
- Filing Date: March 30, 2026
- Status: Pending. A PGR is a trial proceeding conducted at the PTAB to review the patentability of one or more claims in a patent.
Generated 4/30/2026, 4:30:37 PM
Proceedings on file (1)
All PTAB activity →AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.
Current assignee: Krisp Technologies, Inc.
PTAB challenges
AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.
Proceedings overview
U.S. Patent 12,417,756 is currently the subject of one active Post-Grant Review (PGR) proceeding before the Patent Trial and Appeal Board (PTAB). This single pending challenge means the patent's claims are currently under review, and no claims have been definitively invalidated or sustained by the PTAB. For a defendant, this creates a dynamic defensive posture, as the patent's validity is still being contested.
PGR2026-00033 — Krisp Technologies, Inc. v. Sanas.AI Inc.
- Type: Post-Grant Review (PGR)
- Filed: 2026-03-30
- Status: Pending. The proceeding has been filed and is currently awaiting an institution decision from the PTAB.
- Judge panel: The judge panel has not yet been publicly assigned or disclosed in an institution decision, as the case is still in the pre-institution phase.
- Petition grounds: The specific claims challenged, prior art asserted, and statutory bases (§ 102 / § 103 / § 112) for the petition are not publicly available in the provided patent text or readily accessible from a general web search at this early stage of the proceeding. Detailed grounds would typically be found in the filed petition and the institution decision.
- Institution decision: The institution decision has not yet been issued. The PTAB typically has six months from the petition filing date to decide whether to institute a PGR. Therefore, an institution decision for PGR2026-00033 is anticipated around September 30, 2026.
- Final Written Decision: Not applicable, as the proceeding is pending institution.
- Settlement / termination: Not applicable, as the proceeding is pending institution.
- Appeal: Not applicable, as no Final Written Decision has been issued.
- Defensive value: This pending PGR indicates an active challenge to the patent's validity. While no claims have been invalidated yet, the outcome of this proceeding could significantly impact the strength of U.S. Patent 12,417,756. A defendant would closely monitor this case, as a successful PGR could lead to the cancellation of some or all of the patent's claims, weakening any infringement assertions. Conversely, if the claims are sustained, it could harden the patent against future challenges.
Strategic summary
As of May 29, 2026, all claims of U.S. Patent 12,417,756 are currently UNTESTED by a final PTAB decision. There is one active Post-Grant Review (PGR2026-00033) challenging the patent's validity, but it is still in the pre-institution phase. Therefore, no claims have been canceled or sustained by the PTAB.
The estoppel landscape has not yet formed for this patent. Estoppel under § 315(e)(2) only applies to petitioners (and their privies) for grounds raised or reasonably could have been raised after a Final Written Decision has been issued. Since PGR2026-00033 is pending institution, no estoppel implications are present at this time.
Regarding pattern signals, Krisp Technologies, Inc., the petitioner in PGR2026-00033, is also the defendant in the district court litigation Sanas.AI Inc. v. Krisp Technologies, Inc. (Case Number: 3:25-cv-05666). This indicates an aggressive defensive strategy by Krisp Technologies, Inc., utilizing both district court litigation and PTAB proceedings to challenge the patent.
Recommended next steps
For a defendant facing assertion of US Patent 12,417,756, the primary next step is to closely monitor the progress of PGR2026-00033. The key upcoming milestone is the institution decision, which is anticipated around September 30, 2026. This decision will reveal which (if any) claims the PTAB has agreed to review and based on what prior art.
- Monitor PTAB Docket for PGR2026-00033: Regularly check the USPTO PTAB E2E portal for updates on PGR2026-00033. The institution decision is crucial.
- The PTAB case number is PGR2026-00033.
- Review Institution Decision: Once issued, thoroughly review the institution decision. This document will detail the PTAB's reasoning for instituting or denying review, including the specific claims and prior art grounds deemed to have a reasonable likelihood of success. This information will be vital for assessing the patent's vulnerability.
- Stay Apprised of Trial Schedule: If the PGR is instituted, note the statutory one-year trial deadline from institution to Final Written Decision. This will include deadlines for discovery, expert reports, and potential oral hearings.
Generated 5/29/2026, 9:05:19 PM
Ownership chain (2)
Asserters network →Structured records extracted from the assignment-history narrative below. Each entity links to its full ownership-network profile.
2025-01-02 · recorded 2025-01-22 · reel 069963/0080 · Assignment of Assignors Interest
Dura, Piotr; Serebryakov, Maxim; Pfeifenberger, Lukas; Escudero, Alvaro; Jha, Ankita; Zhang, ShawnSANAS.AI INC., CALIFORNIA
Correspondent: · Rutan & Tucker
internal transfer of inventor rights to the original assignee
2025-08-02 · recorded 2025-08-04 · reel 071924/0929 · Assignment of Assignors Interest
NARAYANA, SHARATH KASHAVASANAS.AI INC., CALIFORNIA
Correspondent: · Rutan & Tucker
internal transfer of inventor rights to the original assignee
Assignment history
Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.
Inventors
The inventors named on U.S. Patent 12,417,756 are:
- Ankita Jha
- Lukas Pfeifenberger
- Piotr Dura
- David Braude
- Alvaro Escudero
- Shawn Zhang
- Maxim Serebryakov
- Sharath Kashava Narayana
Based on available information, these inventors are associated with Sanas Ai Inc., with some being co-founders or key team members. Sharath Keshava Narayana is noted as CEO and co-founder, and Shawn Zhang and Maxim Serebryakov are also co-founders. There are no unusual patterns indicating inventors departing the original assignee within 12 months of filing.
Original assignee
The original assignee named on the issued patent is Sanas Ai Inc.
Sanas Ai Inc. is a privately-held software developer that offers a real-time Speech AI platform. Their primary line of business involves providing real-time accent translation, speech enhancement, and language translation technologies, particularly for enterprise and global communication platforms like call centers.
The company actively ships products embodying the claims, specifically their "Accent Translation" technology, which modulates accents in real-time while preserving unique voices and emotions. They have grown significantly, reporting $50M in revenue in January 2026 and supporting over 750,000 users globally.
Sanas Ai Inc. is currently operating and is a Series B company with significant venture backing.
Assignment timeline
The following is a chronological list of every recorded assignment for U.S. Patent 12,417,756, based on the provided patent document's legal events.
2025-01-02 to 2025-01-17 (executed) / recorded 2025-01-22 — Reel 069963/0080
- Conveyance: Assignment of Assignors Interest
- Assignor: Ankita Jha, Lukas Pfeifenberger, Piotr Dura, and others
- Assignee: Sanas.ai Inc., California
- Correspondent: Not available in provided data.
- Context: Initial assignment from multiple inventors to their employer.
2025-08-02 (executed) / recorded 2025-08-04 — Reel 071924/0929
- Conveyance: Assignment of Assignors Interest
- Assignor: Sharath Kashava Narayana
- Assignee: Sanas.ai Inc., California
- Correspondent: Not available in provided data.
- Context: Initial assignment from an inventor to their employer.
There are no further recorded assignments for this patent available in the provided data, meaning Sanas.ai Inc. remains the current assignee.
Timeline diagram
timeline
title Ownership of US 12417756
2025 : Inventors assign to Sanas.ai Inc
: Inventor Narayana assigns to Sanas
NPE / troll-pattern signals
Shell-entity transfer — not present. The patent was assigned directly from the inventors to Sanas.ai Inc., an operating company that develops and markets products embodying the claims. There is no evidence of a transfer to a licensing-only LLC.
Known asserter in the chain — not present. Sanas.ai Inc. is an operating company, not identified as a known NPE. They are currently asserting this patent in a district court case against Krisp Technologies, Inc., which appears to be a competitor.
Repeat correspondent across the chain — unclear. Correspondent information is not available in the provided data for the recorded assignments (Reel 069963/0080, Reel 071924/0929).
Cascading transfers — not present. There are only two assignments, both from inventors to the operating company, and they are not consecutive transfers through chained LLCs.
Pre-litigation transfer — not present. The patent was issued on September 16, 2025, and infringement claims were added to a lawsuit by Sanas.ai Inc. on September 23, 2025. While the assertion happened very close to the issue date, the assignments to Sanas.ai Inc. occurred before the patent issued and well before the lawsuit update. The transfer was from inventors to the operating company, not a transfer to enable litigation.
Bankruptcy fire-sale — not present. Sanas.ai Inc. is an active, well-funded operating company.
Privateering — not present. There is no indication of a transfer to an NPE asserting on behalf of an operating company. Sanas.ai Inc. itself is asserting the patent.
Defensive aggregator (anti-NPE) — not present. The current assignee is Sanas.ai Inc., an operating company.
Verdict
Operating-company assertion
Sanas.ai Inc., the current assignee, is an active operating company that develops and markets products related to real-time accent mimicking. The assignments on record (Reel 069963/0080 and Reel 071924/0929) are initial transfers from the inventors to their employer, which is standard practice. Sanas.ai Inc. is currently asserting this patent against Krisp Technologies, Inc., which appears to be a direct competitor in the speech AI space.
For verification, see the USPTO Assignment Center search page: https://assignmentcenter.uspto.gov/ (search for patent number 12417756).
Generated 5/29/2026, 9:05:33 PM
Prior art
Earlier patents, publications, and products that may anticipate or render the claims unpatentable.
Analysis of Prior Art Cited in U.S. Patent 12,417,756
Based on the patent documentation for U.S. Patent 12,417,756, the following patent documents have been cited as prior art by the examiner. This analysis details the most relevant of these citations and their potential impact on the patent's claims under 35 U.S.C. § 102 (Anticipation).
A patent claim is anticipated if a single prior art reference discloses each and every element of the claim. The independent claims of the '756 patent (claims 1, 7, and 14) broadly cover a system, method, and non-transitory computer-readable medium for modifying a second user's speech to mimic a first user's accent while preserving the second user's natural voice characteristics, using machine learning models.
Key Prior Art and Potential Anticipation
The following references are identified as the most relevant to the claims of U.S. Patent 12,417,756.
1. U.S. Patent 9,129,602 B1: "Mimicking user speech patterns"
- Full Citation: US Patent 9,129,602 B1
- Assignee: Amazon Technologies, Inc.
- Filing Date: December 14, 2012
- Publication Date: September 8, 2015
- Brief Description: This patent describes a system that can generate synthesized speech that mimics the speech patterns of a user. It involves receiving a user's speech, analyzing it to determine speech patterns (like pitch, prosody, and accent), and then using these patterns to generate new speech in the user's voice. The goal is to make synthesized speech, such as from a digital assistant, sound more like the user it is interacting with.
- Potential Anticipation: This reference appears highly relevant. It discloses the core concept of analyzing a user's speech to extract patterns (analogous to the "first user's accent features" and "second user's natural voice") and using those to synthesize new speech.
- Claims 1, 7, and 14: The '602 patent's disclosure of analyzing user speech to determine patterns like accent and prosody and using those to generate new speech could be argued to anticipate the process of extracting accent features and synthesizing a modified output. The key distinction would be whether the '602 patent explicitly teaches the combination of a first user's accent with the preservation of a second user's distinct natural voice. If "mimicking user speech patterns" is interpreted broadly enough to cover this combination, it could potentially anticipate these claims.
2. U.S. Patent Application Publication 2020/0193971 A1: "System and methods for accent and dialect modification"
- Full Citation: US 2020/0193971 A1
- Assignee: i2x GmbH
- Filing Date: December 13, 2018
- Publication Date: June 18, 2020
- Brief Description: This application describes a system for modifying a speaker's accent or dialect in real-time. It involves capturing audio, identifying accent or dialect features, and transforming the speech to a target accent or dialect while aiming to maintain the speaker's voice identity. It is particularly aimed at call centers to help agents be more clearly understood.
- Potential Anticipation: This reference is also highly relevant as it explicitly addresses real-time accent modification while preserving the speaker's voice.
- Claims 1, 7, and 14: The '971 application appears to describe all the key steps of the '756 patent's independent claims. It teaches capturing speech (second user), modifying its accent to a target (first user's accent), and importantly, "maintaining the speaker's voice identity" (preserving the natural voice). The distinction may lie in the specific techniques used for analysis and synthesis (e.g., the '756 patent's specific mention of MFCC or unique fingerprints), but the overall process seems to be disclosed.
3. U.S. Patent 11,134,217 B1: "System that provides video conferencing with accent modification and multiple video overlaying"
- Full Citation: US Patent 11,134,217 B1
- Assignee: Surendra Goel
- Filing Date: January 11, 2021
- Publication Date: September 28, 2021
- Brief Description: This patent details a video conferencing system that includes real-time accent modification. A participant's speech is modified to an accent that is selected or deemed more understandable to other participants. The system is designed to improve clarity in multi-lingual or multi-accent conversations.
- Potential Anticipation: This reference is relevant due to its application of accent modification in a real-time communication context.
- Claims 1, 7, and 14: The '217 patent describes modifying a speaker's accent in real-time within a conferencing system. It inherently involves capturing a second user's speech and modifying it. The central question for anticipation would be whether it explicitly teaches the analysis of a first user's accent to use as the target and the specific step of preserving the second user's "natural voice" characteristics as defined in the '756 patent. If the modification is to a generic "standard" accent rather than mimicking a specific participant, it may not fully anticipate the claims.
4. U.S. Patent Application Publication 2024/0161764 A1: "Accent personalization for speakers and listeners"
- Full Citation: US 2024/0161764 A1
- Assignee: Dell Products L.P.
- Filing Date: November 9, 2022
- Publication Date: May 16, 2024
- Brief Description: This application covers a system that personalizes audio by modifying a speaker's accent to match a listener's preference or native accent. The system can detect a listener's accent and convert the speaker's audio to that accent in real-time to improve comprehension.
- Potential Anticipation: This is a very strong reference, published before the priority date of the '756 patent. It describes the core functionality of the invention.
- Claims 1, 7, and 14: The '764 application teaches modifying a speaker's accent to match a listener's accent, which directly corresponds to the '756 patent's concept of mimicking a "first user" (the listener) for the benefit of a "second user" (the speaker's output). The disclosure of personalizing the accent for the listener strongly implies that the system analyzes the listener's speech characteristics. It likely also discusses preserving the speaker's voice to avoid unnatural output. This reference has a high probability of anticipating the independent claims.
Other Notable Cited References
- US 2023/0267941 A1 ("Personalized Accent and/or Pace of Speaking Modulation for Audio/Video Streams"): Discloses modifying a speaker's accent and pace for a listener, which is conceptually similar.
- US 2024/0146560 A1 ("Participant Audio Stream Modification Within A Conference"): Describes modifying a participant's audio in a conference, which could include accent modification, to improve clarity.
Summary of Analysis
The prior art cited against U.S. Patent 12,417,756, particularly US 9,129,602 B1, US 2020/0193971 A1, and US 2024/0161764 A1, appears to be highly relevant. These documents describe the core concepts of analyzing speech to identify accent and voice characteristics, and using this analysis to modify a speaker's accent in real-time while preserving their voice identity.
The strength of an anticipation argument under 35 U.S.C. § 102 would depend on whether a single one of these references discloses every limitation of the independent claims. Based on the descriptions, the '971 and '764 applications seem to come closest to disclosing the entire claimed process. This body of prior art likely forms the basis for the pending Post-Grant Review (PGR2026-00033) filed by Krisp Technologies, Inc., which challenges the validity of the patent's claims.
Generated 4/30/2026, 4:31:53 PM
Obviousness
Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.
Obviousness Analysis of U.S. Patent 12,417,756 under 35 U.S.C. § 103
This analysis examines whether the claimed invention in U.S. Patent 12,417,756 would have been obvious to a Person Having Ordinary Skill in the Art (PHOSITA) at the time of the invention, based on the prior art cited in the patent's prosecution history. An invention is considered obvious under 35 U.S.C. § 103 if the differences between the claimed invention and the prior art are such that the subject matter as a whole would have been obvious to a PHOSITA.
A PHOSITA in this field would typically possess a graduate degree in computer science or electrical engineering and have several years of experience in speech processing, machine learning, and digital signal processing. This individual would be familiar with technologies such as speech synthesis, voice conversion, speaker identification, and the application of machine learning models to audio data.
The independent claims (1, 7, and 14) of the '756 patent cover a process that can be broken down into the following key steps:
- Analyze a first user's speech using machine learning to extract their specific accent features.
- Analyze a second user's speech to generate characteristics of their distinct "natural voice."
- Synthesize a modified version of the second user's speech that combines the first user's accent with the second user's preserved natural voice.
- Provide this modified speech as audio output in real-time.
Several combinations of the cited prior art references could be argued to render these claims obvious.
Combination 1: US 2024/0161764 A1 (Dell) in view of US 2020/0193971 A1 (i2x GmbH)
This combination presents a strong argument for obviousness.
Primary Reference: US 2024/0161764 A1 ("Accent personalization for speakers and listeners")
The Dell '764 application serves as an excellent base reference. It explicitly teaches the core concept of the '756 patent: modifying a speaker's accent to match a listener's accent in real-time to improve comprehension. In the context of the '756 patent's claims, the Dell "listener" is the "first user," and the Dell "speaker" is the "second user." The '764 application therefore discloses capturing a second user's speech and modifying its accent to mimic that of a first user. A PHOSITA would understand that for this system to be commercially viable, the output must sound natural and not robotic, which implies the preservation of the speaker's core vocal identity.Secondary Reference: US 2020/0193971 A1 ("System and methods for accent and dialect modification")
The i2x '971 application explicitly addresses a known challenge in voice modification: preserving the speaker's unique voice. It teaches a system for transforming speech to a target accent while specifically aiming to "maintain the speaker's voice identity."Motivation to Combine and Obviousness:
A PHOSITA, starting with the system taught by Dell ('764), would be motivated to improve the quality and naturalness of the accent-modified audio output. A common and predictable problem in speech synthesis and conversion is the loss of the original speaker's vocal characteristics, making the output sound artificial. The PHOSITA would have been motivated to look for known techniques to solve this problem. The i2x ('971) application provides an explicit solution by teaching how to maintain the speaker's voice identity during accent modification.Combining these teachings would have been a matter of applying a known technique (preserving speaker identity from i2x) to an existing system (listener-based accent conversion from Dell) to achieve a predictable and desired result (a more natural-sounding accent-modified output). Therefore, it would have been obvious to combine Dell '764 and i2x '971 to arrive at the invention claimed in the '756 patent.
Combination 2: US 2020/0193971 A1 (i2x GmbH) in view of US 2024/0161764 A1 (Dell)
This presents an alternative but equally strong argument.
Primary Reference: US 2020/0193971 A1 (i2x GmbH)
The i2x '971 application teaches a system for modifying a speaker's accent to a "target accent" while preserving the speaker's voice identity. This discloses the majority of the technical steps claimed in the '756 patent: analyzing a second user's speech, preserving its natural voice characteristics, and modifying the accent.Secondary Reference: US 2024/0161764 A1 (Dell)
The missing element in the i2x reference is the specific nature of the "target accent." The i2x system is described in the context of a call center, where the target might be a "standard" or more neutral accent. The Dell '764 application teaches a specific and advantageous application: making the target accent the accent of the listener in a conversation to maximize clarity for that specific listener.Motivation to Combine and Obviousness:
A PHOSITA tasked with improving the i2x ('971) system for real-time communication (e.g., video conferencing, as taught by US 11,134,217 B1) would seek ways to make the accent modification more effective. Instead of converting all speakers to a single, pre-defined target accent, it would have been an obvious and logical improvement to dynamically adapt the target accent to that of the other participant(s) in the conversation. The Dell ('764) reference provides the exact blueprint for this improvement: analyze the listener's accent and use it as the target for conversion.The motivation is to enhance the primary goal of the system—improving communication clarity. By personalizing the accent conversion for the listener, the system becomes more effective. This combination represents the use of a known technique (listener-based targeting from Dell) to improve a known system (accent modification with voice preservation from i2x).
Conclusion
The independent claims of U.S. Patent 12,417,756 appear to be obvious over combinations of the prior art cited during its prosecution. The prior art, particularly the Dell '764 and i2x '971 publications, already established the key inventive concepts: real-time accent conversion to match a listener's accent and the necessity of preserving the speaker's natural voice identity during this process. The combination of these teachings would have been a predictable step for a Person Having Ordinary Skill in the Art seeking to improve the quality and effectiveness of real-time communication systems. The claimed invention is a synthesis of known elements from the prior art, solving a known problem to achieve a predictable result. This high degree of overlap and clear motivation to combine prior art elements forms a strong basis for the pending Post-Grant Review (PGR2026-00033).
Generated 4/30/2026, 4:32:28 PM
Extensions
Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.
Term, Continuation, and Family Data for U.S. Patent 12,417,756
As of April 26, 2026, this analysis details the patent term, application history, and related family members for U.S. Patent 12,417,756, based on the provided authoritative patent documentation.
Patent Term Adjustments (PTA) and Extensions (PTE)
A review of the provided patent text for U.S. Patent 12,417,756 does not indicate that any Patent Term Adjustment (PTA) or Patent Term Extension (PTE) has been granted. The prosecution from the filing date (January 17, 2025) to the issue date (September 16, 2025) was approximately eight months, which is well within the three-year period for which the USPTO may grant PTA for its own delays. There is no information to suggest any other type of delay that would warrant a term adjustment.
Projected Expiration Date
The term of a U.S. patent is typically 20 years from the earliest effective non-provisional filing date. For U.S. Patent 12,417,756, the key dates are:
- Priority Date: August 1, 2024 (from U.S. Provisional Application No. 63/678,180)
- Filing Date: January 17, 2025
The patent term is calculated from the filing date of the non-provisional application, not the provisional application's priority date. Therefore, the projected expiration date is 20 years from the filing date.
- Projected Expiration Date: January 17, 2045
This expiration date is contingent upon the timely payment of all required maintenance fees and assumes the patent is not invalidated or has its term changed for any other reason, such as through the pending Post-Grant Review (PGR2026-00033).
Continuation and Divisional Applications
The documentation for U.S. Patent 12,417,756 indicates the existence of a related continuation application.
- Continuation Application: U.S. Application Number 19/303,881 was filed on August 19, 2025. This application claims priority to the same provisional application as the '756 patent and is listed as a "Related Child Application." This application has been published as US 2026/0038479 A1.
No divisional applications are mentioned in the provided documentation.
Patent Family Members
The patent family for U.S. Patent 12,417,756 consists of the applications that are linked by claiming priority to the same provisional application.
- U.S. Provisional Application No. 63/678,180: Filed August 1, 2024. This is the originating application establishing the priority date for the family.
- U.S. Patent No. 12,417,756: Based on application no. 19/027,799, filed January 17, 2025.
- U.S. Patent Application Publication No. US 2026/0038479 A1: Based on continuation application no. 19/303,881, filed August 19, 2025.
These documents constitute the known U.S. patent family for this invention as of the current date.
Generated 4/30/2026, 4:32:48 PM
Derivative works
Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.
Defensive Disclosure and Prior Art Generation for Real-Time Accent Mimicking
Publication Date: April 30, 2026
Subject Matter: Derivative works and extensions related to the technology disclosed in U.S. Patent 12,417,756. This document is intended to enter the public domain to serve as prior art for future inventions in the field of speech processing and voice modification.
Axis 1: Algorithmic and Architectural Substitution
Derivative 1.1: Adversarial Accent-Style Transfer Network
Enabling Description: This derivative replaces the distinct analysis and synthesis modules with a unified Generative Adversarial Network (GAN) architecture. The system comprises one generator and two discriminator networks. The generator (G) receives the second user's speech waveform (S2) and a target accent embedding vector (E1) extracted from the first user's speech. It outputs a modified waveform (S_mod). The first discriminator (D_accent) is trained to distinguish between S_mod and authentic speech from the first user (S1), forcing G to learn the accent features. The second discriminator (D_identity) is trained to distinguish the speaker identity of S_mod from the original speaker S2, ensuring that G preserves the natural voice characteristics. The loss function for G is a weighted sum of the adversarial losses from both discriminators, ensuring a balance between accent accuracy and speaker preservation.
Mermaid Diagram:
graph TD subgraph User 1 S1[Speech Waveform] --> AE[Accent Encoder] AE --> E1[Accent Embedding Vector] end subgraph User 2 S2[Speech Waveform] --> G[Generator] S2 --> DI[Speaker Identity Encoder] DI --> ID2[Identity Vector] end E1 --> G G --> S_mod[Modified Waveform] subgraph Training / Discrimination S_mod --> D_accent[Accent Discriminator] S1_samples[Real S1 Samples] --> D_accent D_accent --> L_accent[Accent Loss] S_mod --> D_identity[Identity Discriminator] S2_samples[Real S2 Samples] --> D_identity D_identity --> L_identity[Identity Loss] end L_accent --> G L_identity --> G
Derivative 1.2: End-to-End Flow-Based Waveform Generation
Enabling Description: This variation utilizes a non-autoregressive, flow-based deep learning model, analogous to VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech), for direct waveform conversion. The system first extracts linguistic features (phonemes) from the second user's speech using an acoustic model. Simultaneously, a speaker encoder generates a speaker embedding vector. The target accent is represented by a separate accent embedding vector. These three inputs (phonemes, speaker embedding, accent embedding) are fed into a conditional variational autoencoder (VAE) with normalizing flows. The model learns to map the distribution of the second user's speech to the distribution of the first user's accent, conditioned on the linguistic content and speaker identity. The output is a modified waveform generated in a single pass, enabling faster-than-real-time synthesis.
Mermaid Diagram:
sequenceDiagram participant S2 as Second User Speech participant ASR as ASR/Phoneme Extractor participant SE as Speaker Encoder participant AE as Accent Encoder (from User 1) participant VAE as Flow-Based VAE participant Vocoder S2->>ASR: Raw Audio ASR->>VAE: Phoneme Sequence (p) S2->>SE: Raw Audio SE->>VAE: Speaker Embedding (e_spk) AE->>VAE: Accent Embedding (e_acc) VAE->>Vocoder: Latent Representation (z_mod) Note over VAE: P(z|p, e_spk, e_acc) Vocoder->>S2: Modified Waveform
Axis 2: Operational Parameter Expansion
Derivative 2.1: Ultra-Low Latency Mimicking via Predictive Phoneme Framing
Enabling Description: To achieve glass-to-glass latency under 10ms for applications like simultaneous interpretation, this system employs a predictive model. The feature extraction pipeline operates on 20ms audio frames. A lightweight LSTM (Long Short-Term Memory) network, running in parallel, analyzes the linguistic content of the incoming speech and predicts the most likely subsequent phoneme sequence for the next 40-60ms. While the current frame is being converted, the synthesis module pre-computes the acoustic features for the predicted phonemes based on the target accent. When the actual audio frames arrive, the system combines the pre-computed features with the real-time prosodic information (pitch, energy) from the user, drastically reducing the synthesis computation time per frame. This predictive buffering minimizes the perceived delay.
Mermaid Diagram:
graph TD A[Audio Input Stream] --> B{Frame Buffer (20ms)} B --> C[Feature Extraction]; B --> D[Linguistic Analysis]; D --> E[Predictive LSTM]; E --> F[Predicted Phoneme Buffer]; F --> G{Pre-computation Module}; C --> H{Accent Translation}; H --> I[Feature Combination]; G --> I; I --> J[Waveform Synthesis]; J --> K[Audio Output Stream];
Derivative 2.2: Accent Mimicking for Hypersonic and Subsonic Frequencies
Enabling Description: This system is designed for scientific and industrial analysis by applying the concept of "accent" to non-human audio signals. For hypersonic analysis, the system analyzes the acoustic signature of airflow over a vehicle traveling at Mach 5+ to establish a "nominal flight accent." It then analyzes real-time acoustic data from sensors on the vehicle, converting it to mimic the nominal accent. Deviations in the required transformation indicate changes in atmospheric conditions or structural integrity. For subsonic applications, it analyzes ultrasonic vocalizations from rodents in a lab. It establishes a "calm accent" (baseline) and converts real-time vocalizations to this baseline. The acoustic distance of the conversion quantifies the animal's stress level in response to stimuli.
Mermaid Diagram:
stateDiagram-v2 state "Hypersonic Application" as H { [*] --> Baseline: Capture nominal flight acoustic signature Baseline --> Monitoring: Real-time sensor data input Monitoring --> Monitoring: Analyze & transform signature to baseline state "Transformation Delta > Threshold" as Alert { note right of Alert Indicates structural flutter or unexpected turbulence end note } Monitoring --> Alert Alert --> [*] } state "Subsonic (Ultrasonic) Application" as S { [*] --> S_Baseline: Record baseline rodent vocalizations (calm state) S_Baseline --> S_Monitoring: Monitor vocalizations after stimulus S_Monitoring --> S_Monitoring: Convert active vocalizations to calm baseline state "Acoustic Distance High" as Stress { note right of Stress Quantifies stress level based on the degree of required conversion end note } S_Monitoring --> Stress Stress --> [*] }
Axis 3: Cross-Domain Application
Derivative 3.1: Aerospace - ATC Accent Simulation for Pilot Training
Enabling Description: In a flight simulator, a text-to-speech engine generates standard Air Traffic Control (ATC) commands. These serve as the "second user speech" (in its base, accent-neutral form). The system stores a library of accent embeddings from real-world recordings of ATCs in challenging airspaces (e.g., Guangzhou, Mexico City, Lagos). The training scenario selects a target accent ("first user accent"). The accent mimicking system modifies the standard TTS output to realistically replicate the chosen regional accent, including its unique cadence, phonology, and intonation, while preserving the clarity of the base TTS voice ("natural voice"). This exposes student pilots to realistic communication challenges in a safe environment.
Mermaid Diagram:
flowchart LR subgraph Simulator Core A[Training Scenario] --> B{Select Target Airspace}; B --> C[Load ATC Accent Embedding]; A --> D[Generate ATC Command Text]; end subgraph Accent Mimicking System D --> E[Standard TTS Engine]; E --> F[Base Speech Output]; C --> G[Accent/Prosody Modifier]; F --> G; G --> H[Accented Speech Output]; end H --> I[Cockpit Audio System];
Derivative 3.2: AgTech - Pathogenic Beehive Acoustics
Enabling Description: The system is used to detect diseases like Varroa mite infestation in beehives. A high-fidelity microphone records the collective buzzing frequency and pattern of a healthy hive, which is used to create an acoustic embedding for the "healthy hive accent." The system then monitors other hives. The buzzing from a monitored hive ("second user speech") is analyzed. The system modifies this buzzing to mimic the "healthy accent." The parameters of the transformation (e.g., required frequency shift, amplitude modulation) correlate with specific pathogenic stressors. A large transformation magnitude indicates a high probability of infestation, triggering an alert for the beekeeper. The "natural voice" preservation corresponds to maintaining the hive's unique baseline hum, distinguishing it from background noise.
Mermaid Diagram:
sequenceDiagram participant Sensor as Hive Acoustic Sensor participant Analyzer as Accent Analyzer participant Transformer as Accent Transformer participant Dashboard as Beekeeper Dashboard Sensor->>Analyzer: Continuous Buzzing Audio (Hive B) note right of Analyzer: Pre-loaded with "Healthy Hive Accent" embedding (from Hive A) Analyzer->>Transformer: Buzzing Audio + Target Healthy Accent Transformer->>Transformer: Calculate Transformation Parameters Transformer->>Dashboard: Send Health Score (based on transform magnitude) alt Health Score < Threshold Dashboard->>Dashboard: Display "Hive B is Unhealthy" else Dashboard->>Dashboard: Display "Hive B is Healthy" end
Axis 4: Integration with Emerging Tech
Derivative 4.1: IoT and AI for Dynamic Acoustic Ambiance Matching
Enabling Description: In a vehicle or smart home, an array of IoT microphones constantly monitors the ambient conversation. An AI model determines the dominant accent and language of the occupants. When the user interacts with the voice assistant, this system modifies the assistant's standard response voice ("second user") to match the detected ambient accent ("first user"). This integration allows the AI assistant to seamlessly blend into the social environment. If the conversation switches accents (e.g., a new passenger joins the car), the IoT sensors trigger the AI to update the target accent embedding in real-time, ensuring the assistant's voice adapts dynamically.
Mermaid Diagram:
graph TD A[IoT Mic Array] --> B(Real-time Audio Stream); B --> C{AI Ambient Accent Detection}; C --> D[Target Accent Profile]; E[User Query] --> F{Voice Assistant}; F --> G[Standard TTS Response]; D --> H(Accent Mimicking Module); G --> H; H --> I[Adapted TTS Response]; I --> J[Speakers];
Derivative 4.2: Blockchain-Verified "Voice Skins" for the Metaverse
Enabling Description: A voice actor creates a unique vocal identity, including a specific accent, and registers it as a "Voice NFT" on a public blockchain (e.g., Ethereum). The NFT's metadata contains the trained accent embedding vector. A user in the metaverse who purchases or licenses this NFT can apply it to their own voice. When the user speaks ("second user"), the system pulls the accent embedding from the blockchain via a smart contract call. It then modifies the user's voice to mimic the NFT's accent ("first user") while preserving the user's own intonation and emotion ("natural voice"). The blockchain transaction ledger provides an immutable, auditable trail of who is authorized to use the voice skin, preventing digital voice impersonation.
Mermaid Diagram:
classDiagram class User { +walletAddress +speak() } class AccentMimickingSystem { +applyVoiceSkin(audio, nftContractAddress) } class Blockchain { +getAccentEmbedding(nftContractAddress) } class VoiceNFT { <<SmartContract>> +ownerAddress +accentEmbeddingVector } User "1" -- "1" AccentMimickingSystem : Interacts with AccentMimickingSystem "1" -- "1" Blockchain : Queries Blockchain "1" -- "*" VoiceNFT : Manages
Axis 5: The "Inverse" or Failure Mode
Derivative 5.1: Accent Anonymization Filter
Enabling Description: This system operates in an inverse "anonymization" mode. It is designed for applications where accent may introduce bias (e.g., automated job screening, anonymous witness testimony). The system analyzes the user's speech, extracts the accent-specific features (phoneme pronunciation, prosody), and also extracts the core vocal identity features (pitch, timbre, formant structure). It then synthesizes a new speech signal using the user's vocal identity features but replaces the accent-specific features with those from a pre-defined, standardized "neutral" accent model (e.g., a generic newscaster accent). The result is speech that is clearly in the user's voice but stripped of any regional or socio-economic accent markers.
Mermaid Diagram:
flowchart TD A[User Speech Input] --> B{Feature Splitter}; B --> C[Accent Features]; B --> D[Vocal Identity Features]; E[Neutral Accent Model] --> F[Neutral Accent Features]; C --> G{Feature Discard}; D --> H{Speech Synthesizer}; F --> H; H --> I[Anonymized Speech Output];
Derivative 5.2: Graceful Degradation to Phonetic Subtitling
Enabling Description: This is a safe-fail mode for high-noise environments where accent conversion could produce unintelligible artifacts. The system continuously calculates a Signal-to-Noise Ratio (SNR) and a confidence score for its accent analysis. If the SNR drops below a pre-set threshold (e.g., 5dB) or the confidence score is low, the system disables audio synthesis entirely. Instead, it performs a real-time speech-to-text conversion of the user's speech. Crucially, it then uses its accent analysis module not to convert the audio, but to generate a phonetic or dialect-aware subtitle. For example, if it detects a Scottish accent saying "I cannae do it," the subtitle might read:
I cannae [can't] do it, providing the original dialect word and its standard equivalent for maximum clarity.Mermaid Diagram:
stateDiagram-v2 [*] --> Monitoring Monitoring: SNR > 5dB and Confidence > 0.8 Monitoring --> Accent_Conversion: Process Audio Accent_Conversion --> Monitoring: Output modified audio Monitoring --> Phonetic_Subtitling: SNR <= 5dB or Confidence <= 0.8 note right of Phonetic_Subtitling 1. Disable audio synthesis 2. Perform STT 3. Annotate text with phonetic/dialect hints end note Phonetic_Subtitling --> Monitoring: Output enhanced subtitles
Combination Prior Art with Open-Source Standards
Combination with WebRTC and Insertable Streams: A system where the accent mimicking algorithm is compiled to WebAssembly (WASM) and deployed as a JavaScript library. In a peer-to-peer WebRTC video conference, the library uses the Insertable Streams for Media API to intercept the raw audio frames from a user's
MediaStreamTrack. The WASM module performs the accent conversion in-browser, modifying the audio frames before they are passed to the RTCRtpSender for encryption and transmission to the remote peer. This enables client-side, real-time accent mimicking in any modern web application without server-side processing.Combination with the Kaldi Speech Recognition Toolkit: A method for improving the accuracy of accent mimicking by leveraging the detailed acoustic models and forced alignment capabilities of the open-source Kaldi toolkit. The second user's speech is first processed by a Kaldi model to generate a precise, time-aligned phoneme transcription. The accent translation module then uses this alignment to perform a more accurate phoneme-to-phoneme mapping and prosody transfer from the target accent, as it knows the exact start and end time of every sound in the source speech.
Combination with Open-Source Voice Assistant Mycroft: An accent mimicking "skill" for the Mycroft open-source voice assistant. The skill allows a user to configure the assistant's voice personality. The user can have a short conversation with Mycroft ("first user speech") in their own accent. Mycroft's skill extracts the accent features and applies them to its own default TTS voice ("second user speech"). Thereafter, all of Mycroft's responses are delivered in its own voice but mimicking the user's regional accent, creating a personalized and localized user experience.
Generated 4/30/2026, 4:33:57 PM
Keep exploring
Other patents in Audio Technology
- US 6832194Here's a concise summary of US Patent 6832194: US Patent 6832194: Audio recognition peripheral system Title: Audio recognition peripheral system Assignee: Sensory Inc [cite: US6832194B1] Inventors: Forrest S. Mozer, Robert E. Savoie…
- US 10276207US patent 10276207, titled "Virtual wireless multitrack recording system," was issued to Zaxcom Inc. on April 30, 2019. The inventors are Glenn Norman Sanders and Howard Glenn Stark. The patent was filed on August 1, 2016, under…
- US 7711443Here is a concise summary of US patent 7711443: US Patent 7711443: Virtual wireless multitrack recording system Assignee: Zaxcom Inc. Inventors: Glenn Norman Sanders, Howard Glenn Stark Filing Date: 2005-07-14 Issue Date: 2010-05-04…
- US 10798509Here is a concise summary of US Patent 10798509: Title: Wearable electronic device displays a 3D zone from where binaural sound emanates Assignee: Eight Khz LLC (Current Assignee); Individual (Original Assignee) Inventors: Philip Scott…
- US 9226090Here's a concise summary of US Patent 9226090: Title: Sound localization for an electronic call [cite: https://patents.google.com/patent/US9226090/en] Assignee: Eight Khz LLC [cite: https://patents.google.com/patent/US9226090/en]…
- US 8315400US Patent 8,315,400: Method and Device for Acoustic Management Control of Multiple Microphones Title: Method and device for acoustic management control of multiple microphones Assignee: The current assignees are Cases2tech LLC and…
- US 11204736The USPTO search is implicitly covered by the provided patent text, which acts as the authoritative source for the patent details. I have successfully extracted the required information from the provided patent text for the patent details…
- US 11900016US Patent 11900016: Concise Summary Title: Multi-frequency sensing method and apparatus using mobile-clusters Assignee: Zophonos Inc. Inventor: Levaughn Denton Filing Date: December 20, 2021 (Application number US17/555,813) Issue Date…
This patent in court (2)
2 tracked lawsuits name US 12417756.