- Filed
- Mar 27, 2026
- Last modified
- Jul 21, 2026
- Petitioner
- Krisp Technologies, Inc.
- Inventor
- Maxim Serebryakov et al
Invalidity dossier
US 11948550
Real-time accent conversion model
Current assignee: Unified Patents
Added 5/12/2026, 11:37:36 PM
Active provider: Google · gemini-2.5-flash
Auto-generating section 1 of 2: Extensions…
Each section takes ~30-60s with web-search grounding. Keep this tab open — sections will fill in below as they complete.
Patent summary
Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.
US patent 11948550, titled "Real-time accent conversion model", was issued to Sanas Ai Inc. on April 2, 2024. The inventors are Maxim Serebryakov and Shawn Zhang. The application was filed on August 27, 2021.
Abstract:
The patent describes techniques for real-time accent conversion. A computing device receives an indication of a first accent and a second accent, along with speech content in the first accent via a microphone. It uses a first machine-learning algorithm, trained with audio data of the first accent, to derive a non-text linguistic representation of the speech content. Based on this representation, a second machine-learning algorithm (trained with audio data of both the first and second accents) synthesizes new audio data that represents the original speech content but in the second accent. Finally, this synthesized audio data is converted into an audible version of the speech content with the second accent.
Plain-Language Overview of Independent Claims:
Independent Claim 1 (System): This claim describes a computer system with at least one processor and non-transitory computer-readable storage. The system is programmed to:
- Train a first artificial intelligence (AI) model using speech recordings from many different people speaking with a particular "first accent." This training involves breaking down and categorizing individual sound segments (frames) of the recorded speech.
- Use this trained first AI model on live speech received through a microphone. This live speech has the "first accent" and consists of specific speech sounds (phonemes) pronounced in a certain way. The AI converts these sounds into a hidden, non-textual, machine-readable representation.
- Employ a second AI model to create new audio. This second AI model was trained using speech data from both the "first accent" and a "second accent." When creating the new audio, it takes the non-textual representation of a sound from the "first accent" and transforms it into a different non-textual representation that corresponds to how that sound would be pronounced in the "second accent." The key here is that the pronunciation of the speech changes from the first accent to the second.
- Convert this newly generated audio data into a final, audible version of the original speech, which now sounds like it is spoken with the "second accent" and includes the updated pronunciations.
Independent Claim 11 (Non-transitory computer-readable medium): This claim covers a non-transitory computer-readable storage medium (e.g., a hard drive, solid-state drive) that contains software instructions. When a computer's processor runs these instructions, it performs the exact same sequence of steps as described in Claim 1: training a first AI, deriving a non-text linguistic representation from received speech, synthesizing new audio with a second accent by mapping phonemes using a second AI, and finally converting that synthesized audio into an audible version with the second accent.
Independent Claim 19 (Method): This claim outlines a method comprising the identical series of actions as detailed in Claim 1 and Claim 11: training a first machine-learning algorithm with multi-speaker audio of a first accent, applying it to derive a non-text linguistic representation from incoming speech, synthesizing new audio in a second accent using a second machine-learning algorithm by mapping phonemes from the first pronunciation to a different second pronunciation, and then converting this synthesized audio into a finished speech output having the second accent with its characteristic pronunciations.
Litigation Status:
As of the patent's fetch date (May 12, 2026), the patent family is involved in litigation. This includes a PTAB case (IPR2026-00272) that has been filed and is currently pending, as well as a US case filed in the California Northern District Court (case 3:25-cv-05666). The first worldwide family litigation has also been filed. The patent currently holds an "Active" legal status, with an anticipated expiration date of August 27, 2041.
Generated 5/29/2026, 5:55:38 PM
Cases on file (2)
Group view →Specific litigation cases in our database that name US patent 11948550. The free-form analysis below may also discuss cases beyond this list.
- IPR2026-00272Patent Trial and Appeal Board (PTAB)Pending
Defendants: Sanas Ai Inc.
- 3:25-cv-05666California Northern District CourtLitigation
Litigation summary
Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.
US patent 11948550 is involved in the following known litigation:
Case Name/Type: Inter Partes Review (IPR) Proceeding
- Plaintiff(s): Unified Patents
- Defendant(s): Sanas Ai Inc. (as the patent owner)
- Jurisdiction: Patent Trial and Appeal Board (PTAB)
- Case number: IPR2026-00272
- Filing date: Not explicitly stated, but the case is noted as "filed" in 2026.
- Outcome or current status: Pending
Case Name/Type: District Court Litigation
- Plaintiff(s): Not explicitly stated in the provided information.
- Defendant(s): Not explicitly stated in the provided information.
- Jurisdiction: California Northern District Court
- Case number: 3:25-cv-05666
- Filing date: Not explicitly stated.
- Outcome or current status: Litigation
Generated 5/29/2026, 5:55:38 PM
Proceedings on file (1)
All PTAB activity →AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.
Current assignee: Unified Patents
PTAB challenges
AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.
Proceedings overview
There is one AIA trial proceeding on file for US patent 11948550. This proceeding is currently active and pending, meaning no claims have been invalidated or sustained yet. This gives a defendant some uncertainty regarding the patent's enforceability, as an active challenge is underway.
IPR2026-00272 — Krisp Technologies, Inc. v. Sanas Ai Inc
- Type: Inter Partes Review
- Filed: 2026-03-27
- Status: Pending. The petition has been filed, and the PTAB is currently reviewing it to determine whether to institute a trial.
- Judge panel: Not yet publicly available, as the institution decision has not been issued.
- Petition grounds: The specific petition grounds, including which claims are challenged and the prior art cited, are not yet publicly available through standard search methods, as the case is still in its early stages before an institution decision. IPRs typically challenge claims under 35 U.S.C. §§ 102 and/or 103.
- Institution decision: Not yet issued. The statutory deadline for an institution decision is typically six months from the petition's filing date, which would be around 2026-09-27.
- Final Written Decision: Not applicable, as the trial has not been instituted.
- Settlement / termination: Not applicable, as the proceeding is pending institution.
- Appeal: Not applicable, as no Final Written Decision has been issued.
- Defensive value: This proceeding indicates that at least one party (Krisp Technologies, Inc.) believes there are grounds to challenge the patent's validity. While no claims have been invalidated, the existence of a pending IPR creates an ongoing risk for the patent owner and could influence settlement negotiations for a defendant. The outcome of the institution decision and any subsequent trial will significantly impact the defensive posture.
Strategic summary
Currently, all claims of US patent 11948550 are UNTESTED in the context of a final PTAB decision. The sole proceeding, IPR2026-00272, is in its nascent stages, with the PTAB yet to decide on institution. This means no claims have been canceled or confirmed patentable by the PTAB.
The estoppel landscape is undeveloped. Should IPR2026-00272 be instituted and proceed to a Final Written Decision, 35 U.S.C. § 315(e)(2) would bar Krisp Technologies, Inc. (and its privies) from asserting invalidity grounds they raised or reasonably could have raised in the IPR. However, for other potential defendants not in privity with Krisp Technologies, Inc., the full range of prior-art grounds remains available. The presence of Unified Patents in connection with related litigation suggests a potential strategy by defensive aggregators, but the petitioner for this IPR is Krisp Technologies, Inc..
Recommended next steps
- Monitor the status of IPR2026-00272 closely, particularly for the institution decision, which is expected around 2026-09-27. The institution decision will reveal which claims, if any, the PTAB agrees to review.
- If facing an assertion, consider conducting a prior art search independent of the IPR grounds, as all prior art grounds are still available for a new defendant.
- If in discussions with the patent owner, the existence of this pending IPR, even at an early stage, may serve as leverage.
Generated 5/29/2026, 5:55:40 PM
Ownership chain (1)
Asserters network →Structured records extracted from the assignment-history narrative below. Each entity links to its full ownership-network profile.
2021-05-12 · recorded 2021-08-31 · reel 057344/0604 · Assignment
SEREBRYAKOV, MAXIM; ZHANG, SHAWNSanas.ai Inc.
Correspondent: R. SCOTT MCKEEVER · POLSINELLI
Transfer of inventors' interest to their employer, the original assignee
Assignment history
Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.
Inventors
- Maxim Serebryakov (Employer: Sanas Ai Inc. at time of filing, based on original assignee information)
- Shawn Zhang (Employer: Sanas Ai Inc. at time of filing, based on original assignee information)
No unusual patterns observed regarding inventor departures, as both assigned their interest to Sanas.ai Inc.
Original assignee
The original assignee named on the patent is Sanas Ai Inc. The company appears to be an operating company that ships a product embodying the claims, as the patent extensively describes "new software technology" and an "accent-conversion application" that performs real-time accent conversion at low latency. Its primary line of business is real-time accent conversion software utilizing machine-learning models. The current status of Sanas Ai Inc. is "Active".
Assignment timeline
- 2021-05-12 (executed) / recorded 2021-08-31 — Reel 057344/0604
- Conveyance: ASSIGNMENT
- Assignor: SEREBRYAKOV, MAXIM; ZHANG, SHAWN
- Assignee: SANAS.AI INC.
- Correspondent: R. SCOTT MCKEEVER; POLSINELLI PC; 1000 LOUISIANA STREET, SUITE 6400; HOUSTON, TX 77002.
- Context: Transfer of inventors' interest to their employer, the original assignee.
Timeline diagram
timeline
title Ownership of US 11948550
2021 : Inventors assign to Sanas.ai Inc.
2024 : Patent granted
NPE / troll-pattern signals
- Shell-entity transfer — not present. The patent was assigned from the inventors to Sanas.ai Inc., which is described as developing and operating the accent conversion software, indicating it is an operating company.
- Known asserter in the chain — not present. Sanas.ai Inc. is not identified as a known patent asserter or NPE on public lists.
- Repeat correspondent across the chain — not present. Only one assignment event is recorded for this patent, therefore no recurrence of a correspondent within this chain.
- Cascading transfers — not present. Only a single assignment from the inventors to the initial operating company is recorded.
- Pre-litigation transfer — not present. The assignment was executed on 2021-05-12 and recorded on 2021-08-31. The earliest recorded litigation event for this patent family (PTAB case IPR2026-00272, US case 3:25-cv-05666) occurred in 2025/2026, well over six months after the assignment.
- Bankruptcy fire-sale — not present. There is no indication of bankruptcy filings by the assignor or assignee.
- Privateering — not present. There is no evidence suggesting a transfer from an operating company to an NPE for assertion purposes.
- Defensive aggregator (anti-NPE) — not present. The current assignee, Sanas.ai Inc., is not a known defensive aggregator.
Verdict
Operating-company assertion. The sole recorded assignment for US11948550 is a standard transfer of inventor rights to their employer, Sanas.ai Inc., an operating company that appears to develop and commercialize the described accent conversion software. There are no signals indicative of NPE activity in the patent's ownership chain.
USPTO Assignment Center search for US11948550: https://assignmentcenter.uspto.gov/ (Search by patent number "11948550")
Generated 5/29/2026, 5:55:48 PM
Prior art
Earlier patents, publications, and products that may anticipate or render the claims unpatentable.
The search results provide summaries and sometimes claims of the cited patents, which will be helpful for the "brief description" and "potential anticipation" sections.
Let's break down the analysis for each identified relevant prior art.
1. US10163451B2 - Accent translation
- Full Citation: US10163451B2, titled "Accent translation," assigned to Amazon Technologies, Inc.
- Publication/Filing Date: Publication date: 2018-12-25. Priority date: 2016-12-21. This is prior art to US11948550 (priority date 2021-05-06).
- Brief Description: This patent describes an accent translation model that adjusts audio characteristics of input audio from a first accent to resemble those of a second accent. It involves performing voice recognition analysis to identify letters, phonemes, words, and other units of speech and then adjusting audio characteristics for these identified portions. The accent translation is performed on speech captured by audio components, such as a microphone.
- Potential Anticipation (35 U.S.C. § 102):
- Shared elements with US11948550: Both patents deal with receiving speech content from a microphone, identifying a first accent, translating it to a second accent, and outputting the converted speech. Both also mention processing at the phoneme level.
- Distinguishing features of US11948550: US11948550 specifically claims deriving a non-text linguistic representation and synthesizing audio data by mapping non-text linguistic representations of different phonemes for pronunciation change. While US10163451B2 mentions identifying phonemes and adjusting audio characteristics, it doesn't explicitly detail the "non-text linguistic representation" as the core intermediate for conversion or the direct mapping of different phonemes for different pronunciations as defined in US11948550 claims. The description of US10163451B2 implies adjusting audio characteristics based on identified phonemes, which could be interpreted as modifying existing phonemes' acoustic features rather than mapping to different phonemes to reflect pronunciation shifts (e.g., 'th' to 'd' in Indian English vs. SAE).
- Claims potentially anticipated: Claims 1, 11, and 19 of US11948550 describe the overarching system, non-transitory computer-readable medium, and method, respectively. US10163451B2 might anticipate elements of these claims related to general accent translation, receiving speech, and outputting converted speech. However, the specific non-text linguistic representation and mapping of different phonemes for different pronunciations steps, as explicitly defined in US11948550 claims, would likely differentiate it. For example, Claim 1: "...derive a non-text linguistic representation of the set of phonemes associated with a first pronunciation... synthesize... mapping at least a first non-text linguistic representation of a first phoneme... to a second non-text linguistic representation of a second phoneme of an updated set of phonemes associated with a second pronunciation... wherein the first and second phonemes are different phonemes." This specific phoneme mapping for pronunciation difference might not be explicitly present in US10163451B2's abstract.
2. CN111462769A - End-to-end accent conversion method
- Full Citation: CN111462769A, titled "End-to-end accent conversion method," assigned to 深圳市声希科技有限公司 (Shenzhen Voice-X Technology Co., Ltd.).
- Publication/Filing Date: Publication date: 2020-07-28. Priority date: 2020-03-30. This is prior art to US11948550 (priority date 2021-05-06).
- Brief Description: This patent describes an end-to-end accent conversion method. While specific details from the abstract aren't fully available in English, "end-to-end" often implies avoiding intermediate representations like text, which aligns with US11948550's non-text approach. Another document discussing "Non-parallel Accent Transfer based on Fine-grained Controllable Accent Modeling" (a non-patent citation, but highly relevant to the concept) mentions an "end-to-end accent conversion approach" that converts non-native accented into native-accented speech without native reference audio during conversion, using independently trained neural networks including a speaker encoder, a multi-speaker TTS model, an accented ASR model, and a neural vocoder. It aims to model prosodic characteristics, like speaking rate and duration, for more native-sounding output.
- Potential Anticipation (35 U.S.C. § 102):
- Shared elements with US11948550: The "end-to-end accent conversion" directly targets the core concept of US11948550. If this method indeed uses non-text linguistic representations for accent conversion and handles pronunciation changes without a STT-TTS bottleneck, it could be highly anticipatory. The concept of "end-to-end" aligns with US11948550's goal of low latency and preserving nuances, by avoiding STT-TTS. The non-patent citation elaborates that it "is the first model that is able to convert non-native accented into native-accented speech without any guidance from native reference audio during conversion phase" and uses "four independently trained neural networks: a speaker encoder, a multi-speaker TTS model, an accented ASR model and a neural vocoder." The use of a TTS model within the "end-to-end" framework could be a point of distinction, as US11948550 explicitly states its VC engine does "not need to predict and generate output speech as a midpoint for the conversion," functioning "more quickly than alternatives such as a STT-TTS approach" (description, col. 9, lines 40-44).
- Distinguishing features of US11948550: The core distinction for US11948550 lies in its explicit "non-text linguistic representation" and the direct mapping of "different phonemes" (Claim 1) for pronunciation. If CN111462769A or related "end-to-end" approaches still implicitly or explicitly rely on text-based intermediates or don't perform the direct phoneme-to-phoneme mapping for pronunciation change as specified, there could be a distinction. The non-patent citation indicates use of "a multi-speaker TTS model," which might imply an intermediate text-like step, distinguishing it from the specific non-text claims of US11948550.
- Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19, particularly their broader scope on "real-time accent conversion" and the use of machine learning algorithms for deriving linguistic representation and synthesizing audio data. The "end-to-end" aspect could anticipate the low latency and continuous conversion aspects of claims 9, 18, and 20. However, the specific non-text linguistic representation and the mapping of different phonemes as the mechanism for pronunciation change would be key to distinguishing.
3. CN112382267A - Method, apparatus, device and storage medium for converting accents
- Full Citation: CN112382267A, titled "Method, apparatus, device and storage medium for converting accents," assigned to 北京有竹居网络技术有限公司 (Beijing Youzhuju Network Technology Co., Ltd.).
- Publication/Filing Date: Publication date: 2021-02-19. Priority date: 2020-11-13. This is prior art to US11948550 (priority date 2021-05-06).
- Brief Description: Similar to CN111462769A, the title directly indicates "converting accents," making it highly relevant. Without a detailed English abstract, it's hard to ascertain the precise technical approach. However, given its publication date and explicit focus on accent conversion, it is likely to be considered strong prior art.
- Potential Anticipation (35 U.S.C. § 102):
- Shared elements with US11948550: The broad scope of converting accents using a method, apparatus, device, and storage medium is directly analogous to the claims of US11948550.
- Distinguishing features of US11948550: Similar to the discussion for CN111462769A, the precise nature of the "linguistic representation" (text-based or non-text-based) and the specific mechanism for handling pronunciation differences (e.g., direct phoneme mapping) would be critical for differentiation. If it relies on a traditional STT-TTS pipeline or only modifies acoustic features without changing phonemes for pronunciation, US11948550 could distinguish itself.
- Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19 regarding the general concept of accent conversion. Further analysis of its full text would be required to determine anticipation of the specific technical details like "non-text linguistic representation" and explicit "phoneme mapping for different pronunciations."
4. US10614826B2 - System and method for voice-to-voice conversion
- Full Citation: US10614826B2, titled "System and method for voice-to-voice conversion," assigned to Modulate, Inc.
- Publication/Filing Date: Publication date: 2020-04-07. Priority date: 2017-05-24. This is prior art to US11948550 (priority date 2021-05-06).
- Brief Description: This patent describes a system and method for voice-to-voice conversion. Voice conversion often focuses on changing speaker identity (timbre, pitch, etc.) rather than accent (which includes pronunciation changes). The background of US11948550 explicitly distinguishes itself from "voice conversion methods that attempt to adjust the audio characteristics (e.g., pitch, intonation, melody, stress) of a first speaker's voice to more closely resemble the audio characteristics of a second speaker's voice," stating "this type of approach does not account for the different pronunciations of certain sounds that are inherent to a given accent."
- Potential Anticipation (35 U.S.C. § 102):
- Shared elements with US11948550: Both patents generally involve converting aspects of a received voice.
- Distinguishing features of US11948550: The primary distinction for US11948550 is its explicit focus on accent conversion, which includes changing pronunciations through mapping of different phonemes in a non-text linguistic representation. If US10614826B2 primarily focuses on voice identity (timbre, pitch) while preserving the original accent's pronunciation, it would not anticipate US11948550's core innovation regarding accent-specific pronunciation changes. The abstract of US9613620B2 (which is a similar type of voice conversion patent, though not directly cited in US11948550's family citations) states it "may determine a given representation configured to associate the first voice characteristics with the second voice characteristics. The device may provide an output indicative of pronunciations of the one or more speech sounds of the first voice according to the second voice characteristics based on the given representation." This sounds like it might touch on pronunciation, but the emphasis is usually on making the first voice sound like the second speaker, not necessarily changing the underlying accent phonemes for pronunciation differences. Given US11948550's explicit distinguishing of "voice conversion" from "accent conversion," it's likely this patent would not fully anticipate the pronunciation-focused claims of US11948550.
- Claims potentially anticipated: It might anticipate very broad aspects of receiving speech and generating modified speech. However, it's unlikely to anticipate the claims specifying "first accent," "second accent," "non-text linguistic representation," and the specific "mapping of different phonemes for different pronunciations."
I should also consider US20140365216A1 ([Apple Inc.](/litigations/by-plaintiff/Apple%20Inc.)) "System and method for user-specified pronunciation of words for speech synthesis and recognition" and US20150170642A1 (Google Inc.) "Identifying substitute pronunciations." These deal with pronunciation which is a key part of accent conversion.
5. US20140365216A1 - System and method for user-specified pronunciation of words for speech synthesis and recognition
- Full Citation: US20140365216A1, titled "System and method for user-specified pronunciation of words for speech synthesis and recognition," assigned to Apple Inc.
- Publication/Filing Date: Publication date: 2014-12-11. Priority date: 2013-06-07. This is prior art to US11948550 (priority date 2021-05-06).
- Brief Description: This patent focuses on allowing users to specify pronunciations for words for speech synthesis and recognition. This is more about customizing how a system handles specific words or phrases, not necessarily real-time, end-to-end accent conversion for continuous speech, which is the focus of US11948550.
- Potential Anticipation (35 U.S.C. § 102): While it touches on "pronunciation," it seems to be in the context of user-defined rules or dictionaries for speech synthesis/recognition rather than the ML-driven, accent-specific, non-text linguistic mapping for continuous accent conversion. It might anticipate the idea of handling different pronunciations, but not the specific mechanism or the real-time, low-latency, end-to-end accent conversion as claimed by US11948550. It likely falls into the category of "STT-TTS" or rule-based systems that US11948550 distinguishes itself from.
6. US20150170642A1 - Identifying substitute pronunciations
- Full Citation: US20150170642A1, titled "Identifying substitute pronunciations," assigned to Google Inc.
- Publication/Filing Date: Publication date: 2015-06-18. Priority date: 2013-12-17. This is prior art to US11948550 (priority date 2021-05-06).
- Brief Description: This patent describes identifying substitute pronunciations, likely for improving speech recognition or synthesis. Similar to US20140365216A1, this appears to be more about handling variations in pronunciation within a speech system, rather than the real-time conversion of a continuous speech stream from one accent to another using a non-text linguistic representation.
- Potential Anticipation (35 U.S.C. § 102): Again, while "pronunciation" is a common theme, the specific method of US11948550 (non-text linguistic representation, ML models, direct phoneme mapping for accent change, real-time) is likely distinguishable.
Now, let me confirm the exact claims I should cite for potential anticipation for US11948550.
Claim 1 is a system claim, Claim 11 is a non-transitory computer-readable medium claim, and Claim 19 is a method claim. These are the independent claims. Claims 2-10 are dependent on Claim 1, Claims 12-18 on Claim 11, and Claims 20-22 on Claim 19. If an independent claim is anticipated, its dependent claims are also potentially anticipated if they do not add a novel feature that distinguishes from the prior art. Since the prompt asks "which claim(s) it potentially anticipates," I should list the independent claims (1, 11, 19) and any dependent claims if they add specific elements that are also present in the prior art, but for the scope of this task, focusing on the independent claims and their core elements will be sufficient. The request specifically asks "which claim(s) it potentially anticipates," implying I should list the claims of US11948550 that are anticipated by the specific prior art.
The core novelty of US11948550, as stated in its "Overview" and "Detailed Description," lies in:
- Direct accent conversion (changing pronunciation), not just voice characteristics.
- Using non-text linguistic representation to avoid STT-TTS, preserving nuances.
- Achieving real-time, low latency.
- Specifically, mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a different second phoneme to reflect pronunciation change. (Claim 1)
Considering these points, the prior art most likely to anticipate would be those that also implement accent conversion by directly changing pronunciation via a non-text intermediate in real-time.
Based on the summaries:
- US10163451B2 ("Accent translation") seems to involve adjusting audio characteristics based on phonemes, which could encompass pronunciation changes. It doesn't explicitly state "non-text linguistic representation" or "mapping different phonemes," but the general idea of accent translation on a phonetic level is close.
- CN111462769A ("End-to-end accent conversion method") and CN112382267A ("Method... for converting accents") are very strong candidates due to their titles explicitly mentioning "accent conversion" and "end-to-end" (for CN111462769A). The description of the related non-patent literature on "end-to-end accent conversion" (citing Wang et al. (2021) and others) uses PPG (phonetic posteriorgrams) which are non-text linguistic representations, and explicitly mentions converting "non-native-accented into native-accented speech" and modeling "prosodic characteristics." This is very close to what US11948550 claims. The key difference I noted was the mention of a "multi-speaker TTS model" in the non-patent reference, which could imply a text-like intermediate, a point US11948550 specifically avoids.
I will focus on the three strongest candidates: US10163451B2, CN111462769A, and CN112382267A. I will briefly mention US10614826B2 and the Apple/Google patents as less direct, but still relevant to the general field.
Final check on the provided patent document to ensure no internal contradictions. The provided document for US11948550 is the granted patent. I will interpret the citations directly from the "Patent Citations" section.
Let's refine the "potential anticipation" for each of the top three based on the claims of US11948550.
The key elements of Claim 1 (and implicitly 11 and 19) of US11948550 are:
- Training a first ML algorithm with first accent audio data.
- Applying first ML algorithm to received speech to derive a non-text linguistic representation of phonemes associated with a first pronunciation.
- Synthesizing, using a second ML algorithm trained with first and second accent audio data, fourth audio data by mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second phoneme of an updated set of phonemes associated with a second pronunciation that is different from the first pronunciation, wherein the first and second phonemes are different phonemes.
- Converting synthesized audio data into a synthesized version of received speech content having the second accent, comprising the updated set of phonemes associated with the second pronunciation.
- Real-time, low latency (claims 9, 18, 20).
This "mapping different phonemes" to change pronunciation, within a "non-text linguistic representation" and in "real-time," is the specific inventive step.
I need to make sure I state that it potentially anticipates, as a full anticipation analysis would require reviewing the full text of each cited patent and a detailed claim comparison.
Clarification on USPTO database search: The prompt asks to "Search the USPTO database for the specific patent number 11948550. Do not return results for similar numbers." My initial instruction was to use the provided patent text as authoritative, which is already the result of such a search. So I will simply confirm that I'm analyzing the provided patent.Here is an analysis of the most relevant prior art for US Patent 11948550, "Real-time accent conversion model," based on the patent's own citations and a review of publicly available abstracts and summaries. The analysis focuses on publications prior to US11948550's priority date of May 6, 2021.
US Patent 11948550 describes a system and method for real-time accent conversion. Its core inventive features include:
- Receiving speech content with a first accent.
- Deriving a non-text linguistic representation of the speech content using a first machine-learning (ML) algorithm (e.g., an ASR engine) trained with audio data of the first accent.
- Synthesizing audio data representative of the speech content having a second accent, using a second ML algorithm (e.g., a VC engine). This synthesis involves mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second, different phoneme to reflect a change in pronunciation.
- Converting this synthesized audio data into a listenable version of the speech, all performed with very low latency (e.g., 50-700 ms) for real-time communication.
- A key distinction is avoiding a speech-to-text (STT) to text-to-speech (TTS) pipeline, which can lose nuances like prosody and emotion and introduce significant latency.
Based on these differentiating aspects, the most relevant prior art citations are analyzed below:
Most Relevant Prior Art for US11948550
1. US10163451B2
- Full Citation: US10163451B2, "Accent translation," assigned to Amazon Technologies, Inc.
- Publication/Filing Date: Published on December 25, 2018; Priority Date: December 21, 2016. This predates the priority date of US11948550 (May 6, 2021) and is therefore prior art.
- Brief Description: This patent describes a system and method for "accent translation" where an accent translation model adjusts the audio characteristics of input audio from a first accent to more closely resemble those of a second accent. It involves performing voice recognition analysis to identify letters, phonemes, words, and other units of speech and then adjusting audio characteristics for these identified portions. Speech spoken by a user is captured by a microphone and translated from a first accent to a second accent.
- Potential Anticipation (35 U.S.C. § 102):
- This patent broadly anticipates the concept of "accent translation" between a first and second accent, including receiving speech from a microphone and outputting converted speech. It also mentions processing at the phoneme level for adjustment.
- However, US10163451B2 does not explicitly detail the use of a "non-text linguistic representation" as the primary intermediate format, nor does it explicitly claim the specific mechanism of "mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second phoneme that is different from the first pronunciation" to achieve the accent conversion, which is a key distinguishing feature of US11948550.
- Claims potentially anticipated: Elements of claims 1, 11, and 19 of US11948550 related to the general idea of receiving speech in a first accent and converting it to a second accent using machine learning for pronunciation changes might be broadly anticipated. The specific "non-text linguistic representation" and "mapping of different phonemes" for pronunciation differences would be points of distinction for US11948550.
2. CN111462769A
- Full Citation: CN111462769A, "End-to-end accent conversion method," assigned to 深圳市声希科技有限公司 (Shenzhen Voice-X Technology Co., Ltd.).
- Publication/Filing Date: Published on July 28, 2020; Priority Date: March 30, 2020. This predates the priority date of US11948550 (May 6, 2021) and is therefore prior art.
- Brief Description: This patent describes an "end-to-end accent conversion method." While a detailed English abstract for this specific Chinese patent isn't readily available in the search results, related non-patent literature discusses "end-to-end accent conversion" using phonetic posteriorgrams (PPG), a non-text linguistic representation. It aims to convert non-native accented speech to native accented speech, modeling prosodic characteristics, and potentially uses neural networks including a speaker encoder, a multi-speaker TTS model, an accented ASR model, and a neural vocoder.
- Potential Anticipation (35 U.S.C. § 102):
- The "end-to-end accent conversion" directly aligns with the primary objective of US11948550. The use of non-text linguistic representations (like PPGs) for accent transfer, as suggested by related works, is also highly relevant to US11948550's claims. The ability to model prosodic characteristics further supports its relevance in preserving speech nuances, similar to US11948550's objective to avoid loss from STT-TTS.
- A potential point of distinction could be the mention of a "multi-speaker TTS model" in some "end-to-end" frameworks, which US11948550 explicitly states its voice conversion engine avoids by operating on encoded linguistic data rather than generating output speech as a midpoint (like an STT-TTS approach).
- Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19 of US11948550, especially regarding the broad concept of real-time accent conversion, the use of ML algorithms, and the aim for low-latency, continuous conversion (e.g., claims 9, 18, 20). The specific details of the non-text linguistic representation and the explicit mapping of different phonemes for pronunciation changes would need a full textual comparison for a definitive assessment of anticipation.
3. CN112382267A
- Full Citation: CN112382267A, "Method, apparatus, device and storage medium for converting accents," assigned to 北京有竹居网络技术有限公司 (Beijing Youzhuju Network Technology Co., Ltd.).
- Publication/Filing Date: Published on February 19, 2021; Priority Date: November 13, 2020. This predates the priority date of US11948550 (May 6, 2021) and is therefore prior art.
- Brief Description: The title explicitly states a "Method, apparatus, device and storage medium for converting accents," indicating a direct focus on the same problem addressed by US11948550. Without a detailed English abstract, the specific technical approach remains unknown. However, its explicit problem domain makes it highly relevant.
- Potential Anticipation (35 U.S.C. § 102):
- The broad scope of converting accents through a method, apparatus, device, and storage medium is directly analogous to the claims of US11948550.
- Similar to CN111462769A, the precise nature of the "linguistic representation" used (text-based or non-text-based) and the specific mechanism for altering pronunciation (e.g., direct phoneme mapping versus only acoustic feature adjustment) would be critical to determine if it fully anticipates the specific claims of US11948550.
- Claims potentially anticipated: This patent could potentially anticipate claims 1, 11, and 19 of US11948550, which describe the general system, non-transitory computer-readable medium, and method for accent conversion. A full analysis of its technical details would be needed to assess anticipation of the specific elements related to "non-text linguistic representation" and "mapping of different phonemes for different pronunciations."
Other Relevant Prior Art (Less Direct)
US10614826B2: "System and method for voice-to-voice conversion," assigned to Modulate, Inc. (Published April 7, 2020). While it addresses "voice conversion," US11948550 explicitly distinguishes itself from voice conversion methods that only adjust audio characteristics (pitch, intonation) without accounting for pronunciation differences inherent to accents. Therefore, it is less likely to anticipate the specific pronunciation-changing aspects of US11948550.
US20140365216A1: "System and method for user-specified pronunciation of words for speech synthesis and recognition," assigned to Apple Inc. (Published December 11, 2014). This patent focuses on user-defined pronunciation rules for words rather than a comprehensive, real-time, ML-driven accent conversion for continuous speech that dynamically maps phonemes based on learned linguistic representations.
US20150170642A1: "Identifying substitute pronunciations," assigned to Google Inc. (Published June 18, 2015). Similar to the Apple patent, this likely deals with handling pronunciation variations within speech systems for improved recognition or synthesis, rather than real-time, end-to-end accent conversion for continuous speech using non-text linguistic representations to specifically alter pronunciation between accents.
Generated 5/29/2026, 5:56:29 PM
Obviousness
Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.
Obviousness Analysis under 35 U.S.C. § 103 for US Patent 11948550
This analysis identifies combinations of prior art references that would render the claims of US patent 11948550 obvious to a person having ordinary skill in the art (POSITA) at the time of the invention (priority date May 6, 2021). The motivation for combining these references is also explained.
The independent claims of US11948550 (Claims 1, 11, and 19) describe a system, non-transitory computer-readable medium, and method, respectively, for real-time accent conversion. Key features include:
- Deriving a non-text linguistic representation from input speech (first accent).
- Synthesizing output audio (second accent) by mapping a first non-text linguistic representation of a first phoneme to a second non-text linguistic representation of a second, different phoneme.
- Utilizing two machine-learning algorithms (one for deriving the linguistic representation, one for synthesis).
- Real-time operation with low latency (e.g., 50-700 ms).
- Preservation of prosodic features.
Combination of References: Zhao et al. (2019) in view of US10163451B2 (Amazon) and Sajjan & Vijaya (2016)
Primary References:
- Zhao, G., Ding, S., & Gutierrez-Osuna, R. (2019). Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams. (Hereinafter "Zhao (2019)")
- US10163451B2: Accent translation (Amazon Technologies, Inc.) (Hereinafter "Amazon '451")
- Sajjan, S. C., & Vijaya, C. (Mar. 2016). Continuous Speech Recognition of Kannada language using triphone modeling. (Hereinafter "Sajjan (2016)")
Detailed Obviousness Rationale:
A person having ordinary skill in the art (POSITA) in speech processing or machine learning, seeking to improve real-time accent conversion, would have been motivated to combine the teachings of Zhao (2019) with Amazon '451, possibly incorporating well-known ASR techniques described in Sajjan (2016).
Problem Addressed in the Art:
The background of US11948550 explicitly identifies shortcomings in existing accent conversion solutions:
- Voice conversion methods that only adjust audio characteristics (e.g., pitch, intonation) fail to account for pronunciation differences (e.g., "th-stopping" in Indian English to Standard American English).
- Speech-to-text (STT) followed by text-to-speech (TTS) approaches introduce significant latency (up to several seconds) and may lose nuances like prosody and emotion.
How the Combination Renders Claims Obvious:
Preamble (System, processor, non-transitory computer-readable medium): Both Zhao (2019) and Amazon '451 describe computer-implemented systems and methods, implicitly requiring processors and non-transitory computer-readable media for their operation. This element is standard for any modern speech processing technology.
Claim Element 1(a) (Training a first ML algorithm... with multi-speaker first accent data, aligning/classifying frames):
- Zhao (2019) describes "Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams". The generation of Phonetic Posteriorgrams (PPGs) from speech, as a non-text linguistic representation, fundamentally relies on an underlying Automatic Speech Recognition (ASR) model. Training such an ASR model requires "speech content captured from a plurality of different speakers" to ensure robustness and generalization.
- Sajjan (2016) explicitly teaches "Continuous Speech Recognition... using triphone modeling". Aligning and classifying speech frames according to monophone and triphone sounds is a standard and well-known technique in ASR training for developing robust phonetic representations. A POSITA would readily apply these established ASR training methods to train the first machine-learning algorithm used to generate PPGs or other non-text linguistic representations.
Claim Element 1(b) (Applying the first ML algorithm to received speech to derive a non-text linguistic representation):
- Zhao (2019) directly teaches deriving "Phonetic Posteriorgrams" (PPGs) from input speech for accent conversion. PPGs are a clear example of a "non-text linguistic representation" as they represent phonetic probabilities over time without full text transcription, directly addressing the patent's stated advantage over STT-TTS. Input speech would be "received via at least one microphone," a common component of any speech processing system.
Claim Element 1(c) (Synthesizing using a second ML algorithm... by mapping first phoneme to a second, different phoneme):
- Zhao (2019) generally teaches "Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams," which involves a second machine-learning algorithm to convert the non-text linguistic representation (PPGs) into synthesized audio of a target accent. This algorithm would be trained with audio data of both the first and second accents.
- Amazon '451 directly teaches the crucial step of modifying pronunciations for accent conversion. Claim 1 of Amazon '451 states a method comprising "modifying at least some of the set of phonemes to a target set of phonemes based at least in part on the target accent". This explicitly covers "mapping at least a first non-text linguistic representation of a first phoneme... to a second non-text linguistic representation of a second phoneme... that is different from the first pronunciation... wherein the first and second phonemes are different phonemes." For instance, changing the phoneme for "th" to the phoneme for "d" or "t" for an Indian English to SAE conversion, as highlighted in US11948550's background. A POSITA would understand that the "set of phonemes" and "acoustic characteristics" in Amazon '451 constitute a form of linguistic representation, which could be implemented as a non-textual representation like PPGs from Zhao (2019).
Claim Element 1(d) (Converting synthesized audio data into a synthesized version... comprising the updated set of phonemes):
- Both Zhao (2019) and Amazon '451 ultimately produce synthesized speech in the target accent. Zhao's method explicitly involves "Synthesizing Speech" from PPGs. This final step of converting the synthesized audio data (e.g., mel spectrograms as mentioned in US11948550) into an audible waveform using a vocoder or similar component (as described in US11948550) is a well-known process in speech synthesis and is implicitly or explicitly taught by both references as the end goal of accent conversion. The output speech would naturally embody the "updated set of phonemes" resulting from the conversion process.
Motivation for Combination:
A POSITA would have been motivated to combine Zhao (2019) and Amazon '451 to address the identified problems in the art:
- To overcome latency and prosody loss of STT-TTS: Zhao (2019)'s use of Phonetic Posteriorgrams (PPGs) provides a direct speech-to-speech conversion path using a "non-text linguistic representation". This approach is known to offer lower latency and better preservation of prosodic features (like pitch, emotion, and pauses) compared to STT-TTS, directly addressing the issues raised in US11948550's background.
- To address pronunciation differences: Amazon '451 directly provides a solution for modifying phonemes based on the target accent. This explicitly remedies the deficiency of prior voice conversion methods that only adjust acoustic characteristics but fail to alter specific pronunciations, a problem central to US11948550.
- Synergy and Predictability: It would be obvious for a POSITA to integrate the phoneme modification capability of Amazon '451 into the low-latency, non-textual framework of Zhao (2019). The "non-text linguistic representation" (PPGs) from Zhao (2019) provides an ideal intermediate format within which the phoneme modifications taught by Amazon '451 could be implemented, thereby creating a real-time accent conversion system that effectively handles both acoustic and phonetic variations.
Additional Obvious Features:
- Real-time operation (Claim 9): Both PPG-based methods (Zhao 2019) and accent translation (Amazon '451) are developed in contexts where real-time performance is highly desirable for communication applications. Optimizing such systems for "low latency" (e.g., 50-700 ms, or 300 ms as mentioned in US11948550's abstract) is a common design goal in speech processing, and within the ordinary skill of the art for these types of systems.
- Preservation of prosodic features (Claim 10): As noted, PPGs (Zhao 2019) inherently preserve more continuous speech characteristics than discrete text, allowing for better retention of prosody, pitch, and emotion, which are explicitly mentioned as advantages in US11948550.
- Multi-speaker first accent, single speaker second accent (Claims 4, 5): The practice of training ASR or voice conversion models with diverse speakers for input robustness and a single, representative speaker for target voice identity is a well-established technique in speech synthesis and conversion, readily apparent to a POSITA.
Therefore, the combination of Zhao (2019) and Amazon '451, possibly supplemented by Sajjan (2016) for standard ASR training methodologies, renders the claimed invention obvious under 35 U.S.C. § 103.
Generated 5/29/2026, 5:56:26 PM
Extensions
Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.
Derivative works
Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.
Keep exploring
More patents asserted by Unified Patents
- US 10749859A concise summary of US Patent 10,749,859 is as follows: Title: File format and platform for storage and verification of credentials Assignee: Cortex MCP Inc Inventor: Shaunt M. Sarkissian Filing Date: May 24, 2019 Issue Date: August 18…
- US 8224794Here is a concise summary of US Patent 8,224,794. Title: Clearinghouse system, method, and process for inventorying and acquiring infrastructure, monitoring and controlling network performance for enhancement, and providing localized…
- US 7930575US Patent 7930575, titled "Microcontroller for controlling power shutdown process," was filed on September 10, 2007, and issued on April 19, 2011. The inventors are Yukari Suginaka, Toshifumi Hamaguchi, Yoshitaka Kitao, and Shinya…
- US 10735488Here's a concise summary of US patent 10735488: US Patent 10735488: Method of downloading digital content to be rendered Title: Method of downloading digital content to be rendered Assignee: Audio Pod Ip LLC (Current Assignee); Audio Pod…
- US 9512025Here is a concise summary of US Patent 9512025: US Patent 9512025 Title: Methods and apparatuses for reducing heat loss from edge directors Assignee: Corning Inc. Inventors: Ren Hua Chung, Ahdi El-Kahlout, David Scott Franzen, Brendan…
- US 10715806US Patent 10,715,806: Video Transcoding with Metadata Title: Systems, methods, and media for transcoding video data Assignee: Divx LLC Inventors: Ivan Vladimirovich Naletov, Sergey Zurpal Filing Date: March 11, 2019 Issue Date: July 14…
- US 9070374Here's a concise summary of US patent 9070374: Patent Number: US9070374B2 Title: Communication apparatus and condition notification method for notifying a used condition of communication apparatus by using a light-emitting device attached…
- US 11744686Summary of US Patent 11744686: Intraoral Device Title: Intraoral device Current Assignee: Solmetex LLC (though reassignment history also lists Incept Inc., Dryshield, LLC, and security interests by Midcap Financial Trust and Churchill…
Other patents in Software Technology & Computing Systems (T)
- US 9954872Here is a concise summary of US Patent 9954872: US Patent 9954872B2: System and method for identifying unauthorized activities on a computer system using a data structure model Title: System and method for identifying unauthorized…
- US 11789941B2US Patent 11789941B2 is titled "Systems, methods, applications, and user interfaces for providing triggers in a system of record." Assignee: People Center Inc. Inventors: Siddhartha Gunda, Kyle Michael Boston, Daniel Robert Buscaglia…
- US 12032940B2Here's a concise summary of US Patent 12032940B2: Title: Multi-platform application integration and data synchronization Assignee: People Center Inc Inventors: Siddhartha Gunda, Kyle Michael Boston, Daniel Robert Buscaglia, Dilanka Theshan…
- US 11435994B1US Patent 11435994B1, titled "Multi-platform application integration and data synchronization," was issued to People Center Inc. Here is a summary of the patent details: Title: Multi-platform application integration and data…
- US 9215236Here is a concise summary of US Patent 9215236: Title: Secure, policy-based communications security and file sharing across mixed media, mixed-communications modalities and extensible to cloud computing such as SOA [cite: The full patent…
- US 9537900Here's a concise summary of US patent 9537900: US Patent 9537900 Title: Systems and methods for serving application specific policies based on dynamic context Assignee: Avaya Inc. Inventors: Sunil Menon, Shailesh Patel Filing Date…
- US 9693030US patent 9693030, titled "Generating alerts based upon detector outputs," was filed on July 28, 2014, and issued on June 27, 2017. The original assignee was Arris Enterprises LLC, with the current assignee listed as Bison Patent Licensing…
- US 11238344I have analyzed US Patent 11238344 and compiled the requested information. Summary of US Patent 11238344 Title: Artificially intelligent systems, devices, and methods for learning and/or using a device's circumstances for autonomous device…
This patent in court (2)
2 tracked lawsuits name US 11948550.