- Filed
- Mar 18, 2026
- Last modified
- Jun 25, 2026
- Petitioner
- Krisp Technologies, Inc.
- Inventor
- Lukas PFEIFENBERGER et al
Invalidity dossier
US 12131745
System and method for automatic alignment of phonetic content for real-time accent conversion
Current assignee: Sanas Ai Inc
Added 5/12/2026, 11:37:51 PM
Active provider: Google · gemini-2.5-flash
Auto-generating section 1 of 2: Extensions…
Each section takes ~30-60s with web-search grounding. Keep this tab open — sections will fill in below as they complete.
Patent summary
Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.
Here's a concise summary of US patent 12131745:
Title: System and method for automatic alignment of phonetic content for real-time accent conversion
Assignee: Sanas Ai Inc.
Inventors: Lukas PFEIFENBERGER, Shawn Zhang
Filing Date: 2024-06-26
Issue Date: 2024-10-29
Abstract: The disclosed technology relates to methods, accent conversion systems, and non-transitory computer readable media for real-time accent conversion. In some examples, a set of phonetic embedding vectors is obtained for phonetic content representing a source accent and obtained from input audio data. A trained machine learning model is applied to the set of phonetic embedding vectors to generate a set of transformed phonetic embedding vectors corresponding to phonetic characteristics of speech data in a target accent. An alignment is determined by maximizing a cosine distance between the set of phonetic embedding vectors and the set of transformed phonetic embedding vectors. The speech data is then aligned to the phonetic content based on the determined alignment to generate output audio data representing the target accent. The disclosed technology transforms phonetic characteristics of a source accent to match the target accent more closely for efficient and seamless accent conversion in real-time applications.
Plain-language overview of independent claims:
- Claim 1 (System): This claim describes an accent conversion system comprising hardware (audio interface, memory, processors) configured to perform several steps. The system first receives audio data, then generates numerical representations of the original speech sounds (first phonetic embedding vectors) representing a source accent. A trained neural network transforms these into new numerical representations (second phonetic embedding vectors) for a target accent. The system then determines a precise, differentiable alignment by maximizing the cosine distance (a measure of similarity) between the original and transformed phonetic embedding vectors. Finally, it aligns the original speech data to this determined alignment to produce output audio data in the target accent.
- Claim 8 (Method): This claim outlines a method for automatic alignment and real-time accent conversion, implemented by an accent conversion system. It involves obtaining phonetic embedding vectors from input audio data (representing a source accent). A trained machine learning model then generates transformed phonetic embedding vectors corresponding to a target accent. An alignment is determined by maximizing the cosine distance between the original and transformed phonetic embedding vectors. Based on this alignment, the speech data is aligned to generate output audio data representing the target accent.
- Claim 15 (Non-transitory computer-readable medium): This claim covers a non-transitory computer-readable medium (e.g., a hard drive) storing instructions. When executed by at least one processor, these instructions cause the processor to perform steps similar to the method of Claim 8: obtaining first phonetic embedding vectors for a source accent from input audio, applying a trained neural network to generate second phonetic embedding vectors for a target accent, determining an alignment by maximizing the cosine distance between the first and second vectors, and aligning the speech data based on this alignment to generate output audio in the target accent.
CAFC 2026 Dockets:
A search of CAFC 2026 dockets for patent number US12131745 did not yield any specific results for this patent.
Generated 5/29/2026, 5:51:55 PM
Cases on file (1)
Group view →Specific litigation cases in our database that name US patent 12131745. The free-form analysis below may also discuss cases beyond this list.
- 3:25-cv-05666California Northern District CourtActive litigation
Litigation summary
Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.
Here is a list of known litigation involving US Patent 12131745 as of April 26, 2026:
1. District Court Litigation
- Jurisdiction: California Northern District Court
- Case Number: 3:25-cv-05666
- Filing Date: Not explicitly stated in the provided information.
- Plaintiff(s): Not explicitly stated in the provided information.
- Defendant(s): Not explicitly stated in the provided information.
- Outcome or Current Status: Active litigation.
2. PTAB Inter Partes Review (IPR)
- Jurisdiction: Patent Trial and Appeal Board (PTAB)
- Case Number: IPR2026-00274
- Filing Date: Not explicitly stated in the provided information.
- Plaintiff(s) / Petitioner: Not explicitly stated beyond "Unified Patents PTAB Data".
- Defendant(s) / Patent Owner: Not explicitly stated, but the current assignee of the patent is Sanas Ai Inc.
- Outcome or Current Status: Pending.
3. Darts-ip Litigation (Worldwide Family Litigation)
The Google Patents page also indicates "First worldwide family litigation filed" via Darts-ip, linking to family ID 93016828. While this refers to the patent family, it points to a broader litigation context for related patents. Details regarding specific plaintiff(s), defendant(s), jurisdiction, case number, filing date, and outcome or current status are not provided in the readily available patent data beyond the mention of the Darts-ip platform.
Generated 5/29/2026, 5:51:55 PM
Proceedings on file (1)
All PTAB activity →AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.
PTAB challenges
AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.
Proceedings overview
One Inter Partes Review (IPR) proceeding has been filed against US patent 12131745, which is currently pending. This means the patent's claims are actively being challenged, and a defendant should closely monitor the outcome of this proceeding as no claims have been definitively invalidated or sustained by the PTAB yet.
IPR2026-00274 — Krisp Technologies, Inc. v. Sanas Ai Inc.
- Type: Inter Partes Review
- Filed: 2026-03-18
- Status: Pending. The petition has been filed, and the PTAB is currently evaluating whether to institute a trial.
- Judge panel: Information regarding the specific judge panel is not yet publicly available at this early stage of the proceeding.
- Petition grounds: Specific claims challenged, prior art, and statutory bases (§ 102 / § 103 / § 112) are typically detailed in the petition, which is not publicly available via simple search at this preliminary stage. However, IPRs generally focus on § 102 (novelty) and § 103 (obviousness) grounds.
- Institution decision: Not yet issued. The PTAB has a statutory deadline to decide whether to institute the IPR.
- Final Written Decision: Not yet issued. A Final Written Decision would only be issued if the IPR is instituted and proceeds to trial.
- Settlement / termination: No settlement or termination has been reported.
- Appeal: No appeal has been filed as no Final Written Decision has been issued.
- Defensive value: This IPR proceeding represents an active challenge to the patent. For a defendant, this means there is a potential for claims of US12131745 to be invalidated, which could weaken any assertion based on those claims. However, until an institution decision or a Final Written Decision is issued, the full defensive impact remains to be seen.
Strategic summary
Currently, all claims of US patent 12131745 remain untested by a Final Written Decision, as the single IPR proceeding, IPR2026-00274, is in its very early stages and is still pending an institution decision. There are no claims that have been definitively CANCELED or SUSTAINED by the PTAB at this time.
Regarding estoppel, since no institution decision or Final Written Decision has been rendered, the estoppel provisions of § 315(e)(2) are not yet in effect. Krisp Technologies, Inc. (the petitioner) and any parties in privity with them would be barred from raising grounds that were raised or reasonably could have been raised in the IPR, should it proceed to a Final Written Decision. For other potential defendants, prior art grounds remain available for challenge until a trial is instituted and concluded. There is no clear pattern of multiple IPRs from the same petitioner or aggressive PTAB appeals by the patent owner at this point, as this is the first and only reported IPR. Unified Patents is listed as a petitioner in a PTAB case IPR2026-00274, indicating a potential defensive aggregator involvement.
Recommended next steps
The IPR2026-00274 proceeding against US12131745 is pending. The next significant milestone will be the institution decision, which the PTAB typically issues within 6 months of the petition filing date (around September 2026). Parties facing assertion of this patent should monitor the status of IPR2026-00274 closely via the USPTO PTAB E2E system to understand the specific claims challenged and the outcome of the institution decision. The Unified Patents portal also indicates this proceeding is pending.
Generated 5/29/2026, 5:52:02 PM
Ownership chain (1)
Asserters network →Structured records extracted from the assignment-history narrative below. Each entity links to its full ownership-network profile.
2024-06-25 · recorded 2024-08-02 · reel 068162/0809 · Assignment of Assignors Interest
PFEIFENBERGER, LUKAS; ZHANG, SHAWNSanas.ai Inc.
Correspondent: BRENTON A. BUTLER · SANAS.AI INC.
internal reorg
Assignment history
Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.
Inventors
- Lukas PFEIFENBERGER (Sanas Ai Inc)
- Shawn Zhang (Sanas Ai Inc)
No unusual patterns noted; inventors are associated with the original assignee at the time of filing.
Original assignee
Sanas Ai Inc. is the original assignee.
Sanas.ai Inc. operates in the field of real-time accent conversion, suggesting they likely ship a product embodying the claims. Their current status appears to be operating, as indicated by their active legal status and continued involvement in patent applications.
Assignment timeline
- 2024-06-25 (executed) / recorded 2024-08-02 — Reel 068162/0809
- Conveyance: Assignment of Assignors Interest
- Assignor: PFEIFENBERGER, LUKAS; ZHANG, SHAWN
- Assignee: SANAS.AI INC.
- Correspondent: BRENTON A. BUTLER, SANAS.AI INC., 1547 PALACE DRIVE, SAN MATEO, CA 94403
- Context: Internal reorg (assignment from inventors to company)
Timeline diagram
timeline
title Ownership of US 12131745
2024 : Inventors assigned to Sanas.ai Inc.
NPE / troll-pattern signals
- Shell-entity transfer — not present. The assignment is from the inventors to Sanas.ai Inc., which appears to be an operating company.
- Known asserter in the chain — not present. Sanas.ai Inc. is not identified as a known asserter.
- Repeat correspondent across the chain — not present. Only one assignment is recorded for this patent, and therefore no recurrence of a correspondent.
- Cascading transfers — not present. Only one assignment is recorded.
- Pre-litigation transfer — unclear. There is a US case filed in California Northern District Court (3:25-cv-05666) related to this patent family, filed sometime in 2025. The assignment occurred in June 2024 (executed) / August 2024 (recorded). Without a precise filing date for the litigation, it is unclear if this falls within a 6-month window prior to litigation.
- Bankruptcy fire-sale — not present. No indication of bankruptcy.
- Privateering — not present. No evidence of privateering.
- Defensive aggregator (anti-NPE) — not present. The chain does not terminate at a defensive aggregator.
Verdict
Insufficient data. Only one assignment from the inventors to the original assignee, Sanas.ai Inc., is recorded. This record does not provide sufficient information to identify any NPE or patent troll patterns. Further, while litigation exists within the patent family, the exact filing date of the litigation is needed to assess a pre-litigation transfer signal.
USPTO Assignment Center Search for US12131745
Generated 5/29/2026, 5:52:03 PM
Prior art
Earlier patents, publications, and products that may anticipate or render the claims unpatentable.
Most Relevant Prior Art for US Patent 12131745
Here is an analysis of the most relevant prior art cited in US patent 12131745, along with a brief description and potential anticipation under 35 U.S.C. § 102. The focus of the prior art analysis is on patents that address accent conversion, speech modification, or phonetic processing using machine learning, particularly neural networks, and methods for aligning speech components.
1. US20220358903A1 - Real-Time Accent Conversion Model
- Full Citation: US20220358903A1 (Sanas.ai Inc., published November 10, 2022)
- Publication/Filing Date: This application was filed on May 6, 2021.
- Brief Description: This patent application describes techniques for real-time accent conversion using machine-learning algorithms. It involves receiving speech content in a first accent, deriving a linguistic representation, and then synthesizing audio data representing the speech content in a second accent. The conversion is performed with low latency (e.g., 50-700 milliseconds).
- Potential Anticipation (35 U.S.C. § 102): This reference appears to broadly anticipate the concept of real-time accent conversion using machine learning models, as described in independent claims 1, 8, and 15 of US12131745. Specifically, it mentions receiving input speech in a first accent and outputting a synthesized version in a second accent, using machine learning. While US12131745 specifically claims generating phonetic embedding vectors and maximizing cosine distance for differentiable alignment, US20220358903A1 broadly covers the real-time accent conversion model and process. The explicit mention of "an updated set of phonemes associated with the second accent" also suggests phonetic transformation. Therefore, claims 1 (system), 8 (method), and 15 (computer-readable medium) might be anticipated to the extent that they claim a general real-time accent conversion system using machine learning for phonetic transformation. However, the specific mechanism of "jointly maximizing a cosine distance between the first phonetic embedding vectors and the second phonetic embedding vectors" for differentiable alignment is a key distinction that would need further analysis for full anticipation.
2. US10176819B2 - Phonetic posteriorgrams for many-to-one voice conversion
- Full Citation: US10176819B2 (The Chinese University Of Hong Kong, published January 8, 2019)
- Publication/Filing Date: The priority date for this patent is July 11, 2016.
- Brief Description: This patent describes a method for converting speech using phonetic posteriorgrams (PPGs). It involves obtaining target speech, generating a PPG based on acoustic features (potentially using a speaker-independent automatic speech recognition system), and generating a mapping between the PPG and segments of the target speech. The method emphasizes non-parallel training data and aims for voice conversion.
- Potential Anticipation (35 U.S.C. § 102): This patent anticipates aspects of generating numerical representations of speech content (phonetic content, phonetic embedding vectors) and using them for conversion. The PPGs described "capture the posterior probabilities of each phonetic class for each specific time frame of one utterance," which is akin to phonetic embedding vectors capturing phonetic characteristics. It also describes a "many-to-one voice conversion" which implies a source and target. While it doesn't explicitly mention "maximizing cosine distance for differentiable alignment," it does deal with aligning phonetic information. Therefore, elements of claims 1, 8, and 15 that involve generating phonetic representations (first phonetic embedding vectors) from input audio and then performing a conversion to a target accent could be considered anticipated.
3. US10186251B1 - Voice conversion using deep neural network with intermediate voice training
- Full Citation: US10186251B1 (Oben, Inc., published January 22, 2019)
- Publication/Filing Date: The priority date for this patent is August 6, 2015.
- Brief Description: This patent focuses on voice conversion using deep neural networks with intermediate voice training. While the detailed description isn't fully available in the provided snippets, the title suggests the use of deep neural networks for voice conversion, which broadly aligns with the machine learning model used in US12131745.
- Potential Anticipation (35 U.S.C. § 102): Given the limited information, it's difficult to pinpoint exact anticipation. However, the use of "deep neural networks" for "voice conversion" could potentially anticipate the "trained accent conversion neural network" mentioned in claims 1, 8, and 15, especially regarding the general application of neural networks for transforming speech characteristics. Further details on the specific training and conversion mechanisms would be needed for a more precise assessment.
4. US20220122579A1 - End-to-end speech conversion
- Full Citation: US20220122579A1 (Google Llc, published April 21, 2022)
- Publication/Filing Date: The priority date for this patent is February 21, 2019.
- Brief Description: This patent application describes end-to-end speech conversion. While details are not extensively provided in the snippets, "end-to-end speech conversion" implies a system that takes speech as input and produces converted speech as output, likely encompassing accent conversion or similar speech attribute modification.
- Potential Anticipation (35 U.S.C. § 102): The broad concept of "end-to-end speech conversion" could anticipate claims 1, 8, and 15 of US12131745 in terms of the overall goal of converting speech characteristics from a source to a target. However, the specific methodology of "differentiable alignment by jointly maximizing a cosine distance between the first phonetic embedding vectors and the second phonetic embedding vectors" is not explicitly mentioned in the available snippets for this reference and would be a distinguishing feature.
5. US20040148161A1 - Normalization of speech accent
- Full Citation: US20040148161A1 (Das Sharmistha S., published July 29, 2004)
- Publication/Filing Date: The filing date is January 28, 2003.
- Brief Description: This patent describes a system and method for normalizing speech accent to produce substantially unaccented or less-heavily accented speech. It modifies characteristics of input signals representing accented speech to form output signals representing the same speech with less or no accent. The normalization can involve adjusting parameters like voice onset time, vowel duration, and word stop-release time.
- Potential Anticipation (35 U.S.C. § 102): This patent directly addresses accent normalization, which is a form of accent conversion. It anticipates the general idea of converting a source accent to a target accent (implicitly a "normalized" or "unaccented" target). While the described techniques for modifying speech characteristics (e.g., adjusting voice onset time) differ from the phonetic embedding vector approach of US12131745, the fundamental purpose of accent conversion is anticipated. Therefore, the overarching goal of claims 1, 8, and 15 regarding "accent conversion" is anticipated. The specific "differentiable alignment by jointly maximizing a cosine distance" using phonetic embedding vectors would be a distinguishing feature.
6. US20230352001A1 - Voice attribute conversion using speech to speech
- Full Citation: US20230352001A1 (Meaning.Team, Inc., published November 2, 2023)
- Publication/Filing Date: The filing date is April 28, 2022.
- Brief Description: This patent describes a computer-implemented method for near real-time adaptation of voice attributes, including accent, using a speech-to-speech (S2S) machine learning model. It involves feeding source audio content with a source voice attribute into a trained S2S ML model to obtain target audio content with a target voice attribute, where both have the same lexical content and are time-synchronized.
- Potential Anticipation (35 U.S.C. § 102): This reference is highly relevant as it explicitly discusses "accent and/or voice identity of source audio content... adapted to a target accent and/or target voice identity in target audio" using an "S2S ML model" for "near real-time adaptation." The mention of "time-synchronized" content suggests an alignment process. This directly anticipates the core functionality of accent conversion in claims 1, 8, and 15. The use of a machine learning model to transform voice attributes from a source to a target, with the output preserving linguistic content, strongly aligns with the independent claims. The specific differentiator for US121331745 would remain the explicit "differentiable alignment by jointly maximizing a cosine distance between the first phonetic embedding vectors and the second phonetic embedding vectors."
7. US20230335107A1 - Reference-Free Foreign Accent Conversion System and Method
- Full Citation: US20230335107A1 (The Texas A&M University, published October 19, 2023)
- Publication/Filing Date: The priority date for this patent is August 24, 2020.
- Brief Description: This patent describes a reference-free foreign accent conversion system and method. While the snippets don't provide extensive detail, the "reference-free" aspect suggests a system that does not require parallel speech data for training, which could be a different approach compared to systems relying on paired samples.
- Potential Anticipation (35 U.S.C. § 102): The overall concept of a "foreign accent conversion system and method" directly anticipates the general purpose of US12131745. Depending on the details of its methodology, particularly how it handles the alignment and conversion process without explicit reference data, it could potentially anticipate aspects of claims 1, 8, and 15. However, the specific "differentiable alignment by jointly maximizing a cosine distance between the first phonetic embedding vectors and the second phonetic embedding vectors" using explicitly defined phonetic embedding vectors might be a distinguishing feature.
8. US20230223006A1 - Voice conversion method and related device
- Full Citation: US20230223006A1 (Huawei Technologies Co., Ltd., published July 13, 2023)
- Publication/Filing Date: The priority date for this patent is September 21, 2020.
- Brief Description: This patent describes a voice conversion method and related device. Without further details from the snippets, it can be inferred that it relates to transforming voice characteristics from one form to another.
- Potential Anticipation (35 U.S.C. § 102): As with other voice conversion patents, the general concept of "voice conversion" could be considered broadly anticipatory to the extent that it encompasses accent conversion. The specifics of how this patent achieves conversion and alignment would determine its direct relevance to the unique claims of US12131745 regarding phonetic embedding vectors and cosine distance maximization for differentiable alignment.
9. US20210193160A1 - Method and apparatus for voice conversion and storage medium
- Full Citation: US20210193160A1 (Ubtech Robotics Corp Ltd., published June 24, 2021)
- Publication/Filing Date: The priority date for this patent is December 24, 2019.
- Brief Description: This patent describes a method and apparatus for voice conversion and a storage medium. Like several other general voice conversion patents, the available information is limited to the title.
- Potential Anticipation (35 U.S.C. § 102): The general concept of "voice conversion" by an "apparatus" or through a "method" could broadly anticipate the system, method, and computer-readable medium claims (1, 8, 15) of US12131745 concerning the fundamental act of converting speech characteristics. The unique elements of phonetic embedding vectors, cosine distance, and differentiable alignment in US12131745 would need to be considered for detailed analysis.
Generated 5/29/2026, 5:52:16 PM
Obviousness
Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.
Obviousness Analysis of US Patent 12131745 under 35 U.S.C. § 103
This analysis identifies combinations of prior art references that would render the claims of US Patent 12131745 obvious to a person having ordinary skill in the art (POSITA) at the time of the invention (priority date of June 27, 2023). The primary inventive step claimed by US12131745 is the use of a "differentiable alignment by jointly maximizing a cosine distance" between source and transformed phonetic embedding vectors for real-time accent conversion, specifically to overcome the limitations of Dynamic Time Warping (DTW) [Description].
Claims Under Consideration
The independent claims (Claim 1 for a system, Claim 8 for a method, and Claim 15 for a non-transitory computer-readable medium) share the following core elements:
- Obtaining phonetic embedding vectors representing a source accent from input audio data.
- Applying a trained machine learning model (neural network) to generate transformed phonetic embedding vectors representing a target accent.
- Determining an alignment by maximizing a cosine distance between the source and transformed phonetic embedding vectors.
- Aligning the speech data to the phonetic content based on this determined alignment to generate output audio data representing the target accent.
Claim 4 further details the cosine distance determination, including normalizing vectors to a unit norm and computing a dot product, which is a standard method for calculating cosine similarity. [Claim 4]
Identified Prior Art Combinations and Motivation
A person having ordinary skill in the art (POSITA) in the field of real-time accent conversion and deep learning, by 2023, would have been acutely aware of the limitations of conventional alignment techniques like Dynamic Time Warping (DTW). The patent itself explicitly highlights these issues: DTW is "non-differentiable and not providing gradient information," requires "two separate steps," "makes it difficult to train an accent conversion model effectively," and suffers from "non-monotonicity and instability" leading to "alignment errors" and "poor-quality audio signals." [Description] The patent also states that these limitations make it "challenging to optimize current accent conversion systems using gradient-based methods, which are widely used in deep learning models." [Description]
The overarching motivation for a POSITA would therefore be to develop an accent conversion system that is fully differentiable, allows for end-to-end training using gradient-based optimization, and provides a stable, monotonic, and efficient alignment, thereby improving the accuracy, naturalness, and real-time performance of accent conversion.
Combination: US20220358903A1 in view of US20220122579A1 and general knowledge of a POSITA
1. Primary Reference: US20220358903A1 (Sanas.ai Inc.)
This patent, titled "Real-Time Accent Conversion Model," is assigned to the same entity (Sanas.ai Inc.) as US12131745 and represents highly relevant prior art. It would have taught a POSITA the fundamental components of an accent conversion system, including:
- Obtaining input audio data and extracting speech characteristics, which would implicitly involve or lead to the generation of phonetic representations or embeddings.
- Utilizing a model (likely a neural network, given the context of "Real-Time Accent Conversion Model") to process these characteristics to convert a source accent to a target accent.
- Generating output audio data in the target accent.
As a "current accent conversion system" at the time of the instant patent's filing, it would likely have suffered from the DTW limitations that US12131745 explicitly aims to overcome, serving as the baseline system a POSITA would seek to improve. [Description]
2. Secondary Reference: US20220122579A1 (Google Llc.)
This patent, titled "End-to-end speech conversion," teaches systems and methods for determining a mapping between source and target speech signals using a deep neural network, specifically highlighting its "end-to-end" nature. In the field of deep learning, "end-to-end" processing explicitly refers to systems where all components are differentiable and can be trained jointly using gradient-based optimization, directly addressing the limitations of non-differentiable components like DTW. This reference would motivate a POSITA to seek differentiable solutions for alignment within speech conversion systems.
3. General Knowledge of a POSITA:
By 2023, a POSITA would have possessed the following common knowledge:
- Limitations of DTW: The non-differentiability of DTW and its drawbacks for end-to-end training in deep learning architectures were widely recognized problems in sequence-to-sequence tasks, including speech processing. [Description]
- Utility of Phonetic Embedding Vectors: Phonetic embedding vectors are a well-established means to numerically represent the phonetic characteristics of speech, making them suitable for machine learning processing. The patent itself describes them as capturing "important features related to pronunciation, intonation, and other phonetic aspects." [Description]
- Cosine Distance/Similarity as a Differentiable Metric: Cosine distance (or similarity) is a standard, mathematically differentiable metric for quantifying the similarity or dissimilarity between two vectors, particularly effective for high-dimensional embeddings. Calculating cosine distance typically involves normalizing vectors to a unit norm and then computing their dot product, both of which are differentiable operations. [Description] In sequence models, cosine similarity is frequently used within attention mechanisms to determine alignment or relevance between different parts of input and output sequences in a differentiable manner.
Motivation for Combination:
A POSITA, tasked with improving the "Real-Time Accent Conversion Model" taught by US20220358903A1, would be strongly motivated by the widely known problems associated with DTW (non-differentiability, instability, multi-step training) as articulated in US12131745. [Description] Recognizing the benefits of "end-to-end speech conversion" as taught by US20220122579A1, the POSITA would seek to replace the non-differentiable DTW alignment with a differentiable alternative to enable efficient gradient-based optimization of the entire accent conversion pipeline.
Given that phonetic embedding vectors are already being processed (as implied by US20220358903A1) and that cosine distance is a standard and differentiable measure of similarity between vectors, it would have been an obvious engineering choice for a POSITA to adapt cosine distance maximization to create a differentiable alignment for these phonetic embedding vectors. Maximizing cosine distance between unit-normed vectors (equivalent to maximizing their dot product) is a differentiable operation that could be integrated directly into the neural network's loss function, thereby achieving a "differentiable alignment" and allowing for end-to-end training of the accent conversion neural network, precisely addressing the deficiencies of DTW. This approach would lead to the improved performance, stability, and speed (e.g., "about twenty times faster than alignment achieved using dynamic time warping (DTW)") described as advantages of the claimed invention. [Description]
Conclusion on Obviousness
The combination of US20220358903A1 (teaching a real-time accent conversion model), US20220122579A1 (teaching the desirability and means for end-to-end differentiable speech conversion systems), and the general knowledge of a POSITA regarding DTW's limitations and the differentiable properties and applications of cosine distance for vector alignment, would render the claims of US12131745 obvious. A POSITA would have been motivated to combine these elements to overcome the known technical problems of non-differentiable alignment in accent conversion, thereby enabling end-to-end optimization and improving the performance and efficiency of such systems.
Generated 5/29/2026, 5:52:42 PM
Extensions
Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.
Derivative works
Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.
Keep exploring
Other patents in Software Technology & Computing Systems (T)
- US 9954872Here is a concise summary of US Patent 9954872: US Patent 9954872B2: System and method for identifying unauthorized activities on a computer system using a data structure model Title: System and method for identifying unauthorized…
- US 11789941B2US Patent 11789941B2 is titled "Systems, methods, applications, and user interfaces for providing triggers in a system of record." Assignee: People Center Inc. Inventors: Siddhartha Gunda, Kyle Michael Boston, Daniel Robert Buscaglia…
- US 12032940B2Here's a concise summary of US Patent 12032940B2: Title: Multi-platform application integration and data synchronization Assignee: People Center Inc Inventors: Siddhartha Gunda, Kyle Michael Boston, Daniel Robert Buscaglia, Dilanka Theshan…
- US 11435994B1US Patent 11435994B1, titled "Multi-platform application integration and data synchronization," was issued to People Center Inc. Here is a summary of the patent details: Title: Multi-platform application integration and data…
- US 9215236Here is a concise summary of US Patent 9215236: Title: Secure, policy-based communications security and file sharing across mixed media, mixed-communications modalities and extensible to cloud computing such as SOA [cite: The full patent…
- US 9537900Here's a concise summary of US patent 9537900: US Patent 9537900 Title: Systems and methods for serving application specific policies based on dynamic context Assignee: Avaya Inc. Inventors: Sunil Menon, Shailesh Patel Filing Date…
- US 9693030US patent 9693030, titled "Generating alerts based upon detector outputs," was filed on July 28, 2014, and issued on June 27, 2017. The original assignee was Arris Enterprises LLC, with the current assignee listed as Bison Patent Licensing…
- US 11238344I have analyzed US Patent 11238344 and compiled the requested information. Summary of US Patent 11238344 Title: Artificially intelligent systems, devices, and methods for learning and/or using a device's circumstances for autonomous device…
This patent in court (1)
1 tracked lawsuit name US 12131745.