Invalidity dossier

US 11948550

Real-time accent conversion model

Current assignee: Unified Patents

Added 5/12/2026, 11:37:36 PM

At a glanceActive PTAB challenge2 lawsuits on fileasserted by Unified PatentsSoftware Technology & Computing Systems (T)

Active provider: Google · gemini-2.5-flash

Auto-generating section 1 of 2: Extensions

Each section takes ~30-60s with web-search grounding. Keep this tab open — sections will fill in below as they complete.

Patent summary

Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.

✓ Generated

US patent 11948550, titled "Real-time accent conversion model", was issued to Sanas Ai Inc. on April 2, 2024. The inventors are Maxim Serebryakov and Shawn Zhang. The application was filed on August 27, 2021.

Abstract:
The patent describes techniques for real-time accent conversion. A computing device receives an indication of a first accent and a second accent, along with speech content in the first accent via a microphone. It uses a first machine-learning algorithm, trained with audio data of the first accent, to derive a non-text linguistic representation of the speech content. Based on this representation, a second machine-learning algorithm (trained with audio data of both the first and second accents) synthesizes new audio data that represents the original speech content but in the second accent. Finally, this synthesized audio data is converted into an audible version of the speech content with the second accent.

Plain-Language Overview of Independent Claims:

  • Independent Claim 1 (System): This claim describes a computer system with at least one processor and non-transitory computer-readable storage. The system is programmed to:

    1. Train a first artificial intelligence (AI) model using speech recordings from many different people speaking with a particular "first accent." This training involves breaking down and categorizing individual sound segments (frames) of the recorded speech.
    2. Use this trained first AI model on live speech received through a microphone. This live speech has the "first accent" and consists of specific speech sounds (phonemes) pronounced in a certain way. The AI converts these sounds into a hidden, non-textual, machine-readable representation.
    3. Employ a second AI model to create new audio. This second AI model was trained using speech data from both the "first accent" and a "second accent." When creating the new audio, it takes the non-textual representation of a sound from the "first accent" and transforms it into a different non-textual representation that corresponds to how that sound would be pronounced in the "second accent." The key here is that the pronunciation of the speech changes from the first accent to the second.
    4. Convert this newly generated audio data into a final, audible version of the original speech, which now sounds like it is spoken with the "second accent" and includes the updated pronunciations.
  • Independent Claim 11 (Non-transitory computer-readable medium): This claim covers a non-transitory computer-readable storage medium (e.g., a hard drive, solid-state drive) that contains software instructions. When a computer's processor runs these instructions, it performs the exact same sequence of steps as described in Claim 1: training a first AI, deriving a non-text linguistic representation from received speech, synthesizing new audio with a second accent by mapping phonemes using a second AI, and finally converting that synthesized audio into an audible version with the second accent.

  • Independent Claim 19 (Method): This claim outlines a method comprising the identical series of actions as detailed in Claim 1 and Claim 11: training a first machine-learning algorithm with multi-speaker audio of a first accent, applying it to derive a non-text linguistic representation from incoming speech, synthesizing new audio in a second accent using a second machine-learning algorithm by mapping phonemes from the first pronunciation to a different second pronunciation, and then converting this synthesized audio into a finished speech output having the second accent with its characteristic pronunciations.

Litigation Status:
As of the patent's fetch date (May 12, 2026), the patent family is involved in litigation. This includes a PTAB case (IPR2026-00272) that has been filed and is currently pending, as well as a US case filed in the California Northern District Court (case 3:25-cv-05666). The first worldwide family litigation has also been filed. The patent currently holds an "Active" legal status, with an anticipated expiration date of August 27, 2041.

Generated 5/29/2026, 5:55:38 PM

Cases on file (2)

Group view →

Specific litigation cases in our database that name US patent 11948550. The free-form analysis below may also discuss cases beyond this list.

Litigation summary

Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.

✓ Generated

US patent 11948550 is involved in the following known litigation:

  1. Case Name/Type: Inter Partes Review (IPR) Proceeding

  2. Case Name/Type: District Court Litigation

    • Plaintiff(s): Not explicitly stated in the provided information.
    • Defendant(s): Not explicitly stated in the provided information.
    • Jurisdiction: California Northern District Court
    • Case number: 3:25-cv-05666
    • Filing date: Not explicitly stated.
    • Outcome or current status: Litigation

Generated 5/29/2026, 5:55:38 PM

Proceedings on file (1)

All PTAB activity →

AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.

Current assignee: Unified Patents

1 active
Pending
Filed
Mar 27, 2026
Last modified
Jul 21, 2026
Petitioner
Krisp Technologies, Inc.
Inventor
Maxim Serebryakov et al

PTAB challenges

AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.

✓ Generated

Proceedings overview

There is one AIA trial proceeding on file for US patent 11948550. This proceeding is currently active and pending, meaning no claims have been invalidated or sustained yet. This gives a defendant some uncertainty regarding the patent's enforceability, as an active challenge is underway.

IPR2026-00272 — Krisp Technologies, Inc. v. Sanas Ai Inc

  • Type: Inter Partes Review
  • Filed: 2026-03-27
  • Status: Pending. The petition has been filed, and the PTAB is currently reviewing it to determine whether to institute a trial.
  • Judge panel: Not yet publicly available, as the institution decision has not been issued.
  • Petition grounds: The specific petition grounds, including which claims are challenged and the prior art cited, are not yet publicly available through standard search methods, as the case is still in its early stages before an institution decision. IPRs typically challenge claims under 35 U.S.C. §§ 102 and/or 103.
  • Institution decision: Not yet issued. The statutory deadline for an institution decision is typically six months from the petition's filing date, which would be around 2026-09-27.
  • Final Written Decision: Not applicable, as the trial has not been instituted.
  • Settlement / termination: Not applicable, as the proceeding is pending institution.
  • Appeal: Not applicable, as no Final Written Decision has been issued.
  • Defensive value: This proceeding indicates that at least one party (Krisp Technologies, Inc.) believes there are grounds to challenge the patent's validity. While no claims have been invalidated, the existence of a pending IPR creates an ongoing risk for the patent owner and could influence settlement negotiations for a defendant. The outcome of the institution decision and any subsequent trial will significantly impact the defensive posture.

Strategic summary

Currently, all claims of US patent 11948550 are UNTESTED in the context of a final PTAB decision. The sole proceeding, IPR2026-00272, is in its nascent stages, with the PTAB yet to decide on institution. This means no claims have been canceled or confirmed patentable by the PTAB.

The estoppel landscape is undeveloped. Should IPR2026-00272 be instituted and proceed to a Final Written Decision, 35 U.S.C. § 315(e)(2) would bar Krisp Technologies, Inc. (and its privies) from asserting invalidity grounds they raised or reasonably could have raised in the IPR. However, for other potential defendants not in privity with Krisp Technologies, Inc., the full range of prior-art grounds remains available. The presence of Unified Patents in connection with related litigation suggests a potential strategy by defensive aggregators, but the petitioner for this IPR is Krisp Technologies, Inc..

Recommended next steps

  • Monitor the status of IPR2026-00272 closely, particularly for the institution decision, which is expected around 2026-09-27. The institution decision will reveal which claims, if any, the PTAB agrees to review.
  • If facing an assertion, consider conducting a prior art search independent of the IPR grounds, as all prior art grounds are still available for a new defendant.
  • If in discussions with the patent owner, the existence of this pending IPR, even at an early stage, may serve as leverage.

Generated 5/29/2026, 5:55:40 PM

Ownership chain (1)

Asserters network →

Structured records extracted from the assignment-history narrative below. Each entity links to its full ownership-network profile.

  1. 2021-05-12 · recorded 2021-08-31 · reel 057344/0604 · Assignment

    SEREBRYAKOV, MAXIM; ZHANG, SHAWNSanas.ai Inc.

    Correspondent: R. SCOTT MCKEEVER · POLSINELLI

    Transfer of inventors' interest to their employer, the original assignee

Assignment history

Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.

✓ Generated

Inventors

  • Maxim Serebryakov (Employer: Sanas Ai Inc. at time of filing, based on original assignee information)
  • Shawn Zhang (Employer: Sanas Ai Inc. at time of filing, based on original assignee information)

No unusual patterns observed regarding inventor departures, as both assigned their interest to Sanas.ai Inc.

Original assignee

The original assignee named on the patent is Sanas Ai Inc. The company appears to be an operating company that ships a product embodying the claims, as the patent extensively describes "new software technology" and an "accent-conversion application" that performs real-time accent conversion at low latency. Its primary line of business is real-time accent conversion software utilizing machine-learning models. The current status of Sanas Ai Inc. is "Active".

Assignment timeline

  • 2021-05-12 (executed) / recorded 2021-08-31 — Reel 057344/0604
    • Conveyance: ASSIGNMENT
    • Assignor: SEREBRYAKOV, MAXIM; ZHANG, SHAWN
    • Assignee: SANAS.AI INC.
    • Correspondent: R. SCOTT MCKEEVER; POLSINELLI PC; 1000 LOUISIANA STREET, SUITE 6400; HOUSTON, TX 77002.
    • Context: Transfer of inventors' interest to their employer, the original assignee.

Timeline diagram

timeline
    title Ownership of US 11948550
    2021 : Inventors assign to Sanas.ai Inc.
    2024 : Patent granted

NPE / troll-pattern signals

  1. Shell-entity transfernot present. The patent was assigned from the inventors to Sanas.ai Inc., which is described as developing and operating the accent conversion software, indicating it is an operating company.
  2. Known asserter in the chainnot present. Sanas.ai Inc. is not identified as a known patent asserter or NPE on public lists.
  3. Repeat correspondent across the chainnot present. Only one assignment event is recorded for this patent, therefore no recurrence of a correspondent within this chain.
  4. Cascading transfersnot present. Only a single assignment from the inventors to the initial operating company is recorded.
  5. Pre-litigation transfernot present. The assignment was executed on 2021-05-12 and recorded on 2021-08-31. The earliest recorded litigation event for this patent family (PTAB case IPR2026-00272, US case 3:25-cv-05666) occurred in 2025/2026, well over six months after the assignment.
  6. Bankruptcy fire-salenot present. There is no indication of bankruptcy filings by the assignor or assignee.
  7. Privateeringnot present. There is no evidence suggesting a transfer from an operating company to an NPE for assertion purposes.
  8. Defensive aggregator (anti-NPE)not present. The current assignee, Sanas.ai Inc., is not a known defensive aggregator.

Verdict

Operating-company assertion. The sole recorded assignment for US11948550 is a standard transfer of inventor rights to their employer, Sanas.ai Inc., an operating company that appears to develop and commercialize the described accent conversion software. There are no signals indicative of NPE activity in the patent's ownership chain.

USPTO Assignment Center search for US11948550: https://assignmentcenter.uspto.gov/ (Search by patent number "11948550")

Generated 5/29/2026, 5:55:48 PM

Prior art

Earlier patents, publications, and products that may anticipate or render the claims unpatentable.

✓ Generated

The search results provide summaries and sometimes claims of the cited patents, which will be helpful for the "brief description" and "potential anticipation" sections.

Let's break down the analysis for each identified relevant prior art.

1. US10163451B2 - Accent translation

  • Full Citation: US10163451B2, titled "Accent translation," assigned to Amazon Technologies, Inc.
  • Publication/Filing Date: Publication date: 2018-12-25. Priority date: 2016-12-21. This is prior art to US11948550 (priority date 2021-05-06).
  • Brief Description: This patent describes an accent translation model that adjusts audio characteristics of input audio from a first accent to resemble those of a second accent. It involves performing voice recognition analysis to identify letters, phonemes, words, and other units of speech and then adjusting audio characteristics for these identified portions. The accent translation is performed on speech captured by audio components, such as a microphone.
  • Potential Anticipation (35 U.S.C. § 102):
    • Shared elements with US11948550: Both patents deal with receiving speech content from a microphone, identifying a first accent, translating it to a second accent, and outputting the converted speech. Both also mention processing at the phoneme level.
    • Distinguishing features of US11948550: US11948550 specifically claims deriving a non-text linguistic representation and synthesizing audio data by mapping non-text linguistic representations of different phonemes for pronunciation change. While US10163451B2 mentions identifying phonemes and adjusting audio characteristics, it doesn't explicitly detail the "non-text linguistic representation" as the core intermediate for conversion or the direct mapping of different phonemes for different pronunciations as defined in US11948550 claims. The description of US10163451B2 implies adjusting audio characteristics based on identified phonemes, which could be interpreted as modifying existing phonemes' acoustic features rather than mapping to different phonemes to reflect pronunciation shifts (e.g., 'th' to 'd' in Indian English vs. SAE).
    • Claims potentially anticipated: Claims 1, 11, and 19 of US11948550 describe the overarching system, non-transitory computer-readable medium, and method, respectively. US10163451B2 might anticipate elements of these claims related to general accent translation, receiving speech, and outputting converted speech. However, the specific non-text linguistic representation and mapping of different phonemes for different pronunciations steps, as explicitly defined in US11948550 claims, would likely differentiate it. For example, Claim 1: "...derive a non-text linguistic representation of the set of phonemes associated with a first pronunciation... synthesize... mapping at least a first non-text linguistic representation of a first phoneme... to a second non-text linguistic representation of a second phoneme of an updated set of phonemes associated with a second pronunciation... wherein the first and second phonemes are different phonemes." This specific phoneme mapping for pronunciation difference might not be explicitly present in US10163451B2's abstract.

2. CN111462769A - End-to-end accent conversion method

  • Full Citation: CN111462769A, titled "End-to-end accent conversion method," assigned to 深圳市声希科技有限公司 (Shenzhen Voice-X Technology Co., Ltd.).
  • Publication/Filing Date: Publication date: 2020-07-28. Priority date: 2020-03-30. This is prior art to US11948550 (priority date 2021-05-06).
  • Brief Description: This patent describes an end-to-end accent conversion method. While specific details from the abstract aren't fully available in English, "end-to-end" often implies avoiding intermediate representations like text, which aligns with US11948550's non-text approach. Another document discussing "Non-parallel Accent Transfer based on Fine-grained Controllable Accent Modeling" (a non-patent citation, but highly relevant to the concept) mentions an "end-to-end accent conversion approach" that converts non-native accented into native-accented speech without native reference audio during conversion, using independently trained neural networks including a speaker encoder, a multi-speaker TTS model, an accented ASR model, and a neural vocoder. It aims to model prosodic characteristics, like speaking rate and duration, for more native-sounding output.
  • Potential Anticipation (35 U.S.C. § 102):
    • Shared elements with US11948550: The "end-to-end accent conversion" directly targets the core concept of US11948550. If this method indeed uses non-text linguistic representations for accent conversion and handles pronunciation changes without a STT-TTS bottleneck, it could be highly anticipatory. The concept of "end-to-end" aligns with US11948550's goal of low latency and preserving nuances, by avoiding STT-TTS. The non-patent citation elaborates that it "is the first model that is able to convert non-native accented into native-accented speech without any guidance from native reference audio during conversion phase" and uses "four independently trained neural networks: a speaker encoder, a multi-speaker TTS model, an accented ASR model and a neural vocoder." The use of a TTS model within the "end-to-end" framework could be a point of distinction, as US11948550 explicitly states its VC engine does "not need to predict and generate output speech as a midpoint for the conversion," functioning "more quickly than alternatives such as a STT-TTS approach" (description, col. 9, lines 40-44).
    • Distinguishing features of US11948550: The core distinction for US11948550 lies in its explicit "non-text linguistic representation" and the direct mapping of "different phonemes" (Claim 1) for pronunciation. If CN111462769A or related "end-to-end" approaches still implicitly or explicitly rely on text-based intermediates or don't perform the direct phoneme-to-phoneme mapping for pronunciation change as specified, there could be a distinction. The non-patent citation indicates use of "a multi-speaker TTS model," which might imply an intermediate text-like step, distinguishing it from the specific non-text claims of US11948550.
    • Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19, particularly their broader scope on "real-time accent conversion" and the use of machine learning algorithms for deriving linguistic representation and synthesizing audio data. The "end-to-end" aspect could anticipate the low latency and continuous conversion aspects of claims 9, 18, and 20. However, the specific non-text linguistic representation and the mapping of different phonemes as the mechanism for pronunciation change would be key to distinguishing.

3. CN112382267A - Method, apparatus, device and storage medium for converting accents

  • Full Citation: CN112382267A, titled "Method, apparatus, device and storage medium for converting accents," assigned to 北京有竹居网络技术有限公司 (Beijing Youzhuju Network Technology Co., Ltd.).
  • Publication/Filing Date: Publication date: 2021-02-19. Priority date: 2020-11-13. This is prior art to US11948550 (priority date 2021-05-06).
  • Brief Description: Similar to CN111462769A, the title directly indicates "converting accents," making it highly relevant. Without a detailed English abstract, it's hard to ascertain the precise technical approach. However, given its publication date and explicit focus on accent conversion, it is likely to be considered strong prior art.
  • Potential Anticipation (35 U.S.C. § 102):
    • Shared elements with US11948550: The broad scope of converting accents using a method, apparatus, device, and storage medium is directly analogous to the claims of US11948550.
    • Distinguishing features of US11948550: Similar to the discussion for CN111462769A, the precise nature of the "linguistic representation" (text-based or non-text-based) and the specific mechanism for handling pronunciation differences (e.g., direct phoneme mapping) would be critical for differentiation. If it relies on a traditional STT-TTS pipeline or only modifies acoustic features without changing phonemes for pronunciation, US11948550 could distinguish itself.
    • Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19 regarding the general concept of accent conversion. Further analysis of its full text would be required to determine anticipation of the specific technical details like "non-text linguistic representation" and explicit "phoneme mapping for different pronunciations."

4. US10614826B2 - System and method for voice-to-voice conversion

  • Full Citation: US10614826B2, titled "System and method for voice-to-voice conversion," assigned to Modulate, Inc.
  • Publication/Filing Date: Publication date: 2020-04-07. Priority date: 2017-05-24. This is prior art to US11948550 (priority date 2021-05-06).
  • Brief Description: This patent describes a system and method for voice-to-voice conversion. Voice conversion often focuses on changing speaker identity (timbre, pitch, etc.) rather than accent (which includes pronunciation changes). The background of US11948550 explicitly distinguishes itself from "voice conversion methods that attempt to adjust the audio characteristics (e.g., pitch, intonation, melody, stress) of a first speaker's voice to more closely resemble the audio characteristics of a second speaker's voice," stating "this type of approach does not account for the different pronunciations of certain sounds that are inherent to a given accent."
  • Potential Anticipation (35 U.S.C. § 102):
    • Shared elements with US11948550: Both patents generally involve converting aspects of a received voice.
    • Distinguishing features of US11948550: The primary distinction for US11948550 is its explicit focus on accent conversion, which includes changing pronunciations through mapping of different phonemes in a non-text linguistic representation. If US10614826B2 primarily focuses on voice identity (timbre, pitch) while preserving the original accent's pronunciation, it would not anticipate US11948550's core innovation regarding accent-specific pronunciation changes. The abstract of US9613620B2 (which is a similar type of voice conversion patent, though not directly cited in US11948550's family citations) states it "may determine a given representation configured to associate the first voice characteristics with the second voice characteristics. The device may provide an output indicative of pronunciations of the one or more speech sounds of the first voice according to the second voice characteristics based on the given representation." This sounds like it might touch on pronunciation, but the emphasis is usually on making the first voice sound like the second speaker, not necessarily changing the underlying accent phonemes for pronunciation differences. Given US11948550's explicit distinguishing of "voice conversion" from "accent conversion," it's likely this patent would not fully anticipate the pronunciation-focused claims of US11948550.
    • Claims potentially anticipated: It might anticipate very broad aspects of receiving speech and generating modified speech. However, it's unlikely to anticipate the claims specifying "first accent," "second accent," "non-text linguistic representation," and the specific "mapping of different phonemes for different pronunciations."

I should also consider US20140365216A1 ([Apple Inc.](/litigations/by-plaintiff/Apple%20Inc.)) "System and method for user-specified pronunciation of words for speech synthesis and recognition" and US20150170642A1 (Google Inc.) "Identifying substitute pronunciations." These deal with pronunciation which is a key part of accent conversion.

5. US20140365216A1 - System and method for user-specified pronunciation of words for speech synthesis and recognition

  • Full Citation: US20140365216A1, titled "System and method for user-specified pronunciation of words for speech synthesis and recognition," assigned to Apple Inc.
  • Publication/Filing Date: Publication date: 2014-12-11. Priority date: 2013-06-07. This is prior art to US11948550 (priority date 2021-05-06).
  • Brief Description: This patent focuses on allowing users to specify pronunciations for words for speech synthesis and recognition. This is more about customizing how a system handles specific words or phrases, not necessarily real-time, end-to-end accent conversion for continuous speech, which is the focus of US11948550.
  • Potential Anticipation (35 U.S.C. § 102): While it touches on "pronunciation," it seems to be in the context of user-defined rules or dictionaries for speech synthesis/recognition rather than the ML-driven, accent-specific, non-text linguistic mapping for continuous accent conversion. It might anticipate the idea of handling different pronunciations, but not the specific mechanism or the real-time, low-latency, end-to-end accent conversion as claimed by US11948550. It likely falls into the category of "STT-TTS" or rule-based systems that US11948550 distinguishes itself from.

6. US20150170642A1 - Identifying substitute pronunciations

  • Full Citation: US20150170642A1, titled "Identifying substitute pronunciations," assigned to Google Inc.
  • Publication/Filing Date: Publication date: 2015-06-18. Priority date: 2013-12-17. This is prior art to US11948550 (priority date 2021-05-06).
  • Brief Description: This patent describes identifying substitute pronunciations, likely for improving speech recognition or synthesis. Similar to US20140365216A1, this appears to be more about handling variations in pronunciation within a speech system, rather than the real-time conversion of a continuous speech stream from one accent to another using a non-text linguistic representation.
  • Potential Anticipation (35 U.S.C. § 102): Again, while "pronunciation" is a common theme, the specific method of US11948550 (non-text linguistic representation, ML models, direct phoneme mapping for accent change, real-time) is likely distinguishable.

Now, let me confirm the exact claims I should cite for potential anticipation for US11948550.
Claim 1 is a system claim, Claim 11 is a non-transitory computer-readable medium claim, and Claim 19 is a method claim. These are the independent claims. Claims 2-10 are dependent on Claim 1, Claims 12-18 on Claim 11, and Claims 20-22 on Claim 19. If an independent claim is anticipated, its dependent claims are also potentially anticipated if they do not add a novel feature that distinguishes from the prior art. Since the prompt asks "which claim(s) it potentially anticipates," I should list the independent claims (1, 11, 19) and any dependent claims if they add specific elements that are also present in the prior art, but for the scope of this task, focusing on the independent claims and their core elements will be sufficient. The request specifically asks "which claim(s) it potentially anticipates," implying I should list the claims of US11948550 that are anticipated by the specific prior art.

The core novelty of US11948550, as stated in its "Overview" and "Detailed Description," lies in:

  1. Direct accent conversion (changing pronunciation), not just voice characteristics.
  2. Using non-text linguistic representation to avoid STT-TTS, preserving nuances.
  3. Achieving real-time, low latency.
  4. Specifically, mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a different second phoneme to reflect pronunciation change. (Claim 1)

Considering these points, the prior art most likely to anticipate would be those that also implement accent conversion by directly changing pronunciation via a non-text intermediate in real-time.

Based on the summaries:

  • US10163451B2 ("Accent translation") seems to involve adjusting audio characteristics based on phonemes, which could encompass pronunciation changes. It doesn't explicitly state "non-text linguistic representation" or "mapping different phonemes," but the general idea of accent translation on a phonetic level is close.
  • CN111462769A ("End-to-end accent conversion method") and CN112382267A ("Method... for converting accents") are very strong candidates due to their titles explicitly mentioning "accent conversion" and "end-to-end" (for CN111462769A). The description of the related non-patent literature on "end-to-end accent conversion" (citing Wang et al. (2021) and others) uses PPG (phonetic posteriorgrams) which are non-text linguistic representations, and explicitly mentions converting "non-native-accented into native-accented speech" and modeling "prosodic characteristics." This is very close to what US11948550 claims. The key difference I noted was the mention of a "multi-speaker TTS model" in the non-patent reference, which could imply a text-like intermediate, a point US11948550 specifically avoids.

I will focus on the three strongest candidates: US10163451B2, CN111462769A, and CN112382267A. I will briefly mention US10614826B2 and the Apple/Google patents as less direct, but still relevant to the general field.

Final check on the provided patent document to ensure no internal contradictions. The provided document for US11948550 is the granted patent. I will interpret the citations directly from the "Patent Citations" section.

Let's refine the "potential anticipation" for each of the top three based on the claims of US11948550.
The key elements of Claim 1 (and implicitly 11 and 19) of US11948550 are:

  1. Training a first ML algorithm with first accent audio data.
  2. Applying first ML algorithm to received speech to derive a non-text linguistic representation of phonemes associated with a first pronunciation.
  3. Synthesizing, using a second ML algorithm trained with first and second accent audio data, fourth audio data by mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second phoneme of an updated set of phonemes associated with a second pronunciation that is different from the first pronunciation, wherein the first and second phonemes are different phonemes.
  4. Converting synthesized audio data into a synthesized version of received speech content having the second accent, comprising the updated set of phonemes associated with the second pronunciation.
  5. Real-time, low latency (claims 9, 18, 20).

This "mapping different phonemes" to change pronunciation, within a "non-text linguistic representation" and in "real-time," is the specific inventive step.

I need to make sure I state that it potentially anticipates, as a full anticipation analysis would require reviewing the full text of each cited patent and a detailed claim comparison.

Clarification on USPTO database search: The prompt asks to "Search the USPTO database for the specific patent number 11948550. Do not return results for similar numbers." My initial instruction was to use the provided patent text as authoritative, which is already the result of such a search. So I will simply confirm that I'm analyzing the provided patent.Here is an analysis of the most relevant prior art for US Patent 11948550, "Real-time accent conversion model," based on the patent's own citations and a review of publicly available abstracts and summaries. The analysis focuses on publications prior to US11948550's priority date of May 6, 2021.

US Patent 11948550 describes a system and method for real-time accent conversion. Its core inventive features include:

  • Receiving speech content with a first accent.
  • Deriving a non-text linguistic representation of the speech content using a first machine-learning (ML) algorithm (e.g., an ASR engine) trained with audio data of the first accent.
  • Synthesizing audio data representative of the speech content having a second accent, using a second ML algorithm (e.g., a VC engine). This synthesis involves mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second, different phoneme to reflect a change in pronunciation.
  • Converting this synthesized audio data into a listenable version of the speech, all performed with very low latency (e.g., 50-700 ms) for real-time communication.
  • A key distinction is avoiding a speech-to-text (STT) to text-to-speech (TTS) pipeline, which can lose nuances like prosody and emotion and introduce significant latency.

Based on these differentiating aspects, the most relevant prior art citations are analyzed below:


Most Relevant Prior Art for US11948550

1. US10163451B2

  • Full Citation: US10163451B2, "Accent translation," assigned to Amazon Technologies, Inc.
  • Publication/Filing Date: Published on December 25, 2018; Priority Date: December 21, 2016. This predates the priority date of US11948550 (May 6, 2021) and is therefore prior art.
  • Brief Description: This patent describes a system and method for "accent translation" where an accent translation model adjusts the audio characteristics of input audio from a first accent to more closely resemble those of a second accent. It involves performing voice recognition analysis to identify letters, phonemes, words, and other units of speech and then adjusting audio characteristics for these identified portions. Speech spoken by a user is captured by a microphone and translated from a first accent to a second accent.
  • Potential Anticipation (35 U.S.C. § 102):
    • This patent broadly anticipates the concept of "accent translation" between a first and second accent, including receiving speech from a microphone and outputting converted speech. It also mentions processing at the phoneme level for adjustment.
    • However, US10163451B2 does not explicitly detail the use of a "non-text linguistic representation" as the primary intermediate format, nor does it explicitly claim the specific mechanism of "mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second phoneme that is different from the first pronunciation" to achieve the accent conversion, which is a key distinguishing feature of US11948550.
    • Claims potentially anticipated: Elements of claims 1, 11, and 19 of US11948550 related to the general idea of receiving speech in a first accent and converting it to a second accent using machine learning for pronunciation changes might be broadly anticipated. The specific "non-text linguistic representation" and "mapping of different phonemes" for pronunciation differences would be points of distinction for US11948550.

2. CN111462769A

  • Full Citation: CN111462769A, "End-to-end accent conversion method," assigned to 深圳市声希科技有限公司 (Shenzhen Voice-X Technology Co., Ltd.).
  • Publication/Filing Date: Published on July 28, 2020; Priority Date: March 30, 2020. This predates the priority date of US11948550 (May 6, 2021) and is therefore prior art.
  • Brief Description: This patent describes an "end-to-end accent conversion method." While a detailed English abstract for this specific Chinese patent isn't readily available in the search results, related non-patent literature discusses "end-to-end accent conversion" using phonetic posteriorgrams (PPG), a non-text linguistic representation. It aims to convert non-native accented speech to native accented speech, modeling prosodic characteristics, and potentially uses neural networks including a speaker encoder, a multi-speaker TTS model, an accented ASR model, and a neural vocoder.
  • Potential Anticipation (35 U.S.C. § 102):
    • The "end-to-end accent conversion" directly aligns with the primary objective of US11948550. The use of non-text linguistic representations (like PPGs) for accent transfer, as suggested by related works, is also highly relevant to US11948550's claims. The ability to model prosodic characteristics further supports its relevance in preserving speech nuances, similar to US11948550's objective to avoid loss from STT-TTS.
    • A potential point of distinction could be the mention of a "multi-speaker TTS model" in some "end-to-end" frameworks, which US11948550 explicitly states its voice conversion engine avoids by operating on encoded linguistic data rather than generating output speech as a midpoint (like an STT-TTS approach).
    • Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19 of US11948550, especially regarding the broad concept of real-time accent conversion, the use of ML algorithms, and the aim for low-latency, continuous conversion (e.g., claims 9, 18, 20). The specific details of the non-text linguistic representation and the explicit mapping of different phonemes for pronunciation changes would need a full textual comparison for a definitive assessment of anticipation.

3. CN112382267A

  • Full Citation: CN112382267A, "Method, apparatus, device and storage medium for converting accents," assigned to 北京有竹居网络技术有限公司 (Beijing Youzhuju Network Technology Co., Ltd.).
  • Publication/Filing Date: Published on February 19, 2021; Priority Date: November 13, 2020. This predates the priority date of US11948550 (May 6, 2021) and is therefore prior art.
  • Brief Description: The title explicitly states a "Method, apparatus, device and storage medium for converting accents," indicating a direct focus on the same problem addressed by US11948550. Without a detailed English abstract, the specific technical approach remains unknown. However, its explicit problem domain makes it highly relevant.
  • Potential Anticipation (35 U.S.C. § 102):
    • The broad scope of converting accents through a method, apparatus, device, and storage medium is directly analogous to the claims of US11948550.
    • Similar to CN111462769A, the precise nature of the "linguistic representation" used (text-based or non-text-based) and the specific mechanism for altering pronunciation (e.g., direct phoneme mapping versus only acoustic feature adjustment) would be critical to determine if it fully anticipates the specific claims of US11948550.
    • Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19 of US11948550, which describe the general system, non-transitory computer-readable medium, and method for accent conversion. A full analysis of its technical details would be needed to assess anticipation of the specific elements related to "non-text linguistic representation" and "mapping of different phonemes for different pronunciations."

Other Relevant Prior Art (Less Direct)

  • US10614826B2: "System and method for voice-to-voice conversion," assigned to Modulate, Inc. (Published April 7, 2020). While it addresses "voice conversion," US11948550 explicitly distinguishes itself from voice conversion methods that only adjust audio characteristics (pitch, intonation) without accounting for pronunciation differences inherent to accents. Therefore, it is less likely to anticipate the specific pronunciation-changing aspects of US11948550.

  • US20140365216A1: "System and method for user-specified pronunciation of words for speech synthesis and recognition," assigned to Apple Inc. (Published December 11, 2014). This patent focuses on user-defined pronunciation rules for words rather than a comprehensive, real-time, ML-driven accent conversion for continuous speech that dynamically maps phonemes based on learned linguistic representations.

  • US20150170642A1: "Identifying substitute pronunciations," assigned to Google Inc. (Published June 18, 2015). Similar to the Apple patent, this likely deals with handling pronunciation variations within speech systems for improved recognition or synthesis, rather than real-time, end-to-end accent conversion for continuous speech using non-text linguistic representations to specifically alter pronunciation between accents.

Generated 5/29/2026, 5:56:29 PM

Obviousness

Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.

✓ Generated

Obviousness Analysis under 35 U.S.C. § 103 for US Patent 11948550

This analysis identifies combinations of prior art references that would render the claims of US patent 11948550 obvious to a person having ordinary skill in the art (POSITA) at the time of the invention (priority date May 6, 2021). The motivation for combining these references is also explained.

The independent claims of US11948550 (Claims 1, 11, and 19) describe a system, non-transitory computer-readable medium, and method, respectively, for real-time accent conversion. Key features include:

  1. Deriving a non-text linguistic representation from input speech (first accent).
  2. Synthesizing output audio (second accent) by mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second, different phoneme.
  3. Utilizing two machine-learning algorithms (one for deriving the linguistic representation, one for synthesis).
  4. Real-time operation with low latency (e.g., 50-700 ms).
  5. Preservation of prosodic features.

Combination of References: Zhao et al. (2019) in view of US10163451B2 (Amazon) and Sajjan & Vijaya (2016)

Primary References:

  • Zhao, G., Ding, S., & Gutierrez-Osuna, R. (2019). Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams. (Hereinafter "Zhao (2019)")
  • US10163451B2: Accent translation (Amazon Technologies, Inc.) (Hereinafter "Amazon '451")
  • Sajjan, S. C., & Vijaya, C. (Mar. 2016). Continuous Speech Recognition of Kannada language using triphone modeling. (Hereinafter "Sajjan (2016)")

Detailed Obviousness Rationale:

A person having ordinary skill in the art (POSITA) in speech processing or machine learning, seeking to improve real-time accent conversion, would have been motivated to combine the teachings of Zhao (2019) with Amazon '451, possibly incorporating well-known ASR techniques described in Sajjan (2016).

Problem Addressed in the Art:
The background of US11948550 explicitly identifies shortcomings in existing accent conversion solutions:

  1. Voice conversion methods that only adjust audio characteristics (e.g., pitch, intonation) fail to account for pronunciation differences (e.g., "th-stopping" in Indian English to Standard American English).
  2. Speech-to-text (STT) followed by text-to-speech (TTS) approaches introduce significant latency (up to several seconds) and may lose nuances like prosody and emotion.

How the Combination Renders Claims Obvious:

  1. Preamble (System, processor, non-transitory computer-readable medium): Both Zhao (2019) and Amazon '451 describe computer-implemented systems and methods, implicitly requiring processors and non-transitory computer-readable media for their operation. This element is standard for any modern speech processing technology.

  2. Claim Element 1(a) (Training a first ML algorithm... with multi-speaker first accent data, aligning/classifying frames):

    • Zhao (2019) describes "Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams". The generation of Phonetic Posteriorgrams (PPGs) from speech, as a non-text linguistic representation, fundamentally relies on an underlying Automatic Speech Recognition (ASR) model. Training such an ASR model requires "speech content captured from a plurality of different speakers" to ensure robustness and generalization.
    • Sajjan (2016) explicitly teaches "Continuous Speech Recognition... using triphone modeling". Aligning and classifying speech frames according to monophone and triphone sounds is a standard and well-known technique in ASR training for developing robust phonetic representations. A POSITA would readily apply these established ASR training methods to train the first machine-learning algorithm used to generate PPGs or other non-text linguistic representations.
  3. Claim Element 1(b) (Applying the first ML algorithm to received speech to derive a non-text linguistic representation):

    • Zhao (2019) directly teaches deriving "Phonetic Posteriorgrams" (PPGs) from input speech for accent conversion. PPGs are a clear example of a "non-text linguistic representation" as they represent phonetic probabilities over time without full text transcription, directly addressing the patent's stated advantage over STT-TTS. Input speech would be "received via at least one microphone," a common component of any speech processing system.
  4. Claim Element 1(c) (Synthesizing using a second ML algorithm... by mapping first phoneme to a second, different phoneme):

    • Zhao (2019) generally teaches "Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams," which involves a second machine-learning algorithm to convert the non-text linguistic representation (PPGs) into synthesized audio of a target accent. This algorithm would be trained with audio data of both the first and second accents.
    • Amazon '451 directly teaches the crucial step of modifying pronunciations for accent conversion. Claim 1 of Amazon '451 states a method comprising "modifying at least some of the set of phonemes to a target set of phonemes based at least in part on the target accent". This explicitly covers "mapping at least a first non-text linguistic representation of a first phoneme... to a second non-text linguistic representation of a second phoneme... that is different from the first pronunciation... wherein the first and second phonemes are different phonemes." For instance, changing the phoneme for "th" to the phoneme for "d" or "t" for an Indian English to SAE conversion, as highlighted in US11948550's background. A POSITA would understand that the "set of phonemes" and "acoustic characteristics" in Amazon '451 constitute a form of linguistic representation, which could be implemented as a non-textual representation like PPGs from Zhao (2019).
  5. Claim Element 1(d) (Converting synthesized audio data into a synthesized version... comprising the updated set of phonemes):

    • Both Zhao (2019) and Amazon '451 ultimately produce synthesized speech in the target accent. Zhao's method explicitly involves "Synthesizing Speech" from PPGs. This final step of converting the synthesized audio data (e.g., mel spectrograms as mentioned in US11948550) into an audible waveform using a vocoder or similar component (as described in US11948550) is a well-known process in speech synthesis and is implicitly or explicitly taught by both references as the end goal of accent conversion. The output speech would naturally embody the "updated set of phonemes" resulting from the conversion process.

Motivation for Combination:

A POSITA would have been motivated to combine Zhao (2019) and Amazon '451 to address the identified problems in the art:

  • To overcome latency and prosody loss of STT-TTS: Zhao (2019)'s use of Phonetic Posteriorgrams (PPGs) provides a direct speech-to-speech conversion path using a "non-text linguistic representation". This approach is known to offer lower latency and better preservation of prosodic features (like pitch, emotion, and pauses) compared to STT-TTS, directly addressing the issues raised in US11948550's background.
  • To address pronunciation differences: Amazon '451 directly provides a solution for modifying phonemes based on the target accent. This explicitly remedies the deficiency of prior voice conversion methods that only adjust acoustic characteristics but fail to alter specific pronunciations, a problem central to US11948550.
  • Synergy and Predictability: It would be obvious for a POSITA to integrate the phoneme modification capability of Amazon '451 into the low-latency, non-textual framework of Zhao (2019). The "non-text linguistic representation" (PPGs) from Zhao (2019) provides an ideal intermediate format within which the phoneme modifications taught by Amazon '451 could be implemented, thereby creating a real-time accent conversion system that effectively handles both acoustic and phonetic variations.

Additional Obvious Features:

  • Real-time operation (Claim 9): Both PPG-based methods (Zhao 2019) and accent translation (Amazon '451) are developed in contexts where real-time performance is highly desirable for communication applications. Optimizing such systems for "low latency" (e.g., 50-700 ms, or 300 ms as mentioned in US11948550's abstract) is a common design goal in speech processing, and within the ordinary skill of the art for these types of systems.
  • Preservation of prosodic features (Claim 10): As noted, PPGs (Zhao 2019) inherently preserve more continuous speech characteristics than discrete text, allowing for better retention of prosody, pitch, and emotion, which are explicitly mentioned as advantages in US11948550.
  • Multi-speaker first accent, single speaker second accent (Claims 4, 5): The practice of training ASR or voice conversion models with diverse speakers for input robustness and a single, representative speaker for target voice identity is a well-established technique in speech synthesis and conversion, readily apparent to a POSITA.

Therefore, the combination of Zhao (2019) and Amazon '451, possibly supplemented by Sajjan (2016) for standard ASR training methodologies, renders the claimed invention obvious under 35 U.S.C. § 103.

Generated 5/29/2026, 5:56:26 PM

Extensions

Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.

Not generated yet. Click Generate to call the active LLM provider with the configured prompt.

Derivative works

Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.

Not generated yet. Click Generate to call the active LLM provider with the configured prompt.

Keep exploring

More patents asserted by Unified Patents

Other patents in Software Technology & Computing Systems (T)

See all Software Technology & Computing Systems (T) patents →

This patent in court (2)

2 tracked lawsuits name US 11948550.