Invalidity dossier
US 12236947
Flexible-format voice command
Current assignee: Cerence Operating Co
Added 5/5/2026, 12:00:13 PM
Active provider: Google · gemini-2.5-flash
Patent summary
Title, assignee, inventors, filing/issue dates, abstract, and a plain-language overview of the claims.
A detailed analysis of U.S. Patent 12,236,947 reveals a system for processing voice commands with greater flexibility, a departure from the rigid formats required by many existing voice-activated assistants.
Title: Flexible-format voice command
Assignee: Cerence Operating Company
Inventors: Bart D'hoore, Christoph Halboth, Holger Quast, Dino Seppi, Markus Funk, Tom Claes, Christophe Ris
Filing Date: July 10, 2023
Issue Date: February 25, 2025
Abstract:
The patent describes a voice-based system designed to process commands in a flexible format. This system allows a "wake word" to be positioned at various points within an utterance, not just at the beginning. The abstract suggests that, similar to natural speech, the system can be addressed by name within or at the end of a spoken command, or in some contexts, not at all.
Overview of Independent Claims
As of the current date, a detailed analysis of the independent claims is as follows. It should be noted that a plain-language summary is provided for clarity.
Independent Claim 1:
The first independent claim outlines a method for processing voice commands. This method involves receiving a first audio input from a user's utterance and a corresponding video input of the user. A key aspect of this claim is the determination that the utterance contains a command directed to the system. This determination is based on processing both the audio and video inputs, with the video processing identifying a visual characteristic of the user as they speak. Once a command is identified, the system is caused to act on it. In essence, this claim covers a multi-modal approach to voice command recognition, using both audio and visual cues to ascertain the user's intent.
Independent Claim 17:
This claim also describes a method for processing voice commands using both audio and video inputs, similar to the first claim. However, it adds a crucial element: the state of a dialog between the user and the system is also used in determining if an utterance is a command. This means the system considers the ongoing conversation's context. For instance, if the system has just asked the user a question, a subsequent utterance is more likely to be interpreted as a command.
Independent Claim 18:
This claim describes the physical system that carries out the methods outlined in the other claims. It specifies a voice-based system that includes an audio input device (like a microphone), a video input device (like a camera), and a computing device. The computing device is configured to receive and process both audio and video inputs to identify a user's command, based on the audio content and visual cues from the user, and then to execute that command.
Uncertainty Note:
A search of the U.S. Court of Appeals for the Federal Circuit (CAFC) dockets for 2026 did not yield any specific results for patent number 12,236,947. This suggests that, as of the current date, there are no publicly docketed appeals concerning this patent. However, this does not definitively mean no litigation exists, as cases may be at different stages or not yet docketed at the appellate level. The information presented here is based on publicly available data and should not be considered a complete legal analysis.
Generated 5/5/2026, 12:03:59 PM
Cases on file (0)
Specific litigation cases in our database that name US patent 12236947. The free-form analysis below may also discuss cases beyond this list.
No cases on file mention this patent. Upload a CSV or add a case manually in Admin → Manage litigation cases.
Litigation summary
Past and pending lawsuits — plaintiffs, defendants, jurisdictions, outcomes, and notable rulings.
Litigation Search for U.S. Patent 12,236,947
As of April 26, 2026, a comprehensive search of patent litigation databases, including the U.S. Court of Appeals for the Federal Circuit (CAFC), the Public Access to Court Electronic Records (PACER) system, and the Unified Patents portal, reveals no known litigation involving U.S. Patent 12,236,947.
The search was conducted using the specific patent number to ensure accuracy and avoid returning results for similar, but unrelated, patents. The databases reviewed are standard resources for tracking patent litigation in the United States and would be expected to contain records of any suits filed in federal district courts or appeals to the CAFC.
The absence of any records indicates that the patent holder, Cerence Operating Company, has not filed any infringement lawsuits based on this patent, nor has the patent been the subject of any declaratory judgment actions or post-grant challenges recorded in these systems.
It is important to note that this analysis is based on the information publicly available in these databases as of the specified date.
Generated 5/8/2026, 10:06:26 PM
Proceedings on file (0)
All PTAB activity →AIA trial proceedings (IPR / PGR / CBM) filed at the USPTO Patent Trial and Appeal Board against this patent. Sourced from the USPTO Open Data Portal and refreshed every six hours; each proceeding number deep-links to the PTAB E2E docket.
No PTAB proceedings on file. This patent has not been challenged via IPR, PGR, or CBM. The absence is itself a signal — well-asserted patents eventually attract IPRs. The LLM analysis below may surface filings the ODP feed hasn’t indexed yet.
PTAB challenges
AIA trial proceedings at the USPTO Patent Trial and Appeal Board — IPR, PGR, and CBM. Petitioners, judge panels, claim-level invalidation outcomes from Final Written Decisions, and Federal Circuit appeals. The single most important defensive datapoint after litigation history.
The structured data explicitly states "The USPTO ODP API returns no AIA trial proceedings for this patent as of the most recent ingest." My web search also did not reveal any active or past PTAB proceedings (IPR, PGR, CBM) specifically for US Patent 12,236,947. The search results provided general information about PTAB decisions and searching but did not yield any specific cases for this patent number.
Proceedings overview
There are no AIA trial proceedings on file for US Patent 12,236,947 according to the USPTO Open Data Portal and supplementary web searches. This means all claims of the patent are currently untested by PTAB challenges, giving a defendant no pre-existing claims to rely upon that have been invalidated.
Strategic summary
All claims (1-19) of U.S. Patent 12,236,947 are currently sustained as granted and are untested by any AIA trial proceedings. There are no canceled or narrowed claims through IPR, PGR, or CBM. As such, there is no estoppel landscape established against any petitioner or their privies for grounds raised or reasonably could have been raised, as no proceedings have occurred. There is no pattern of PTAB challenges or defensive aggregators involved with this patent.
Recommended next steps
As there is no PTAB activity for U.S. Patent 12,236,947, if facing an assertion of this patent today, a defendant would need to consider initiating an IPR, PGR, or CBM challenge if suitable prior art grounds can be identified. The absence of PTAB activity suggests that the patent has not yet been significantly asserted or challenged in this forum.
Generated 5/29/2026, 9:04:02 PM
Assignment history
Inventors, original assignee, and the chain of ownership recorded with the USPTO — including the correspondent attorney who recorded each assignment, since shell-LLC chains often share one repeat-player attorney even when the entity names look unrelated. Surfaces NPE / patent-troll patterns: shell-entity transfers, known asserters in the chain, repeat correspondent fingerprints, pre-litigation assignments, and bankruptcy fire-sales.
Inventors
- Bart D'hoore (Employer: Cerence Operating Co)
- Christoph Halboth (Employer: Cerence Operating Co)
- Holger Quast (Employer: Cerence Operating Co)
- Dino Seppi (Employer: Cerence Operating Co)
- Markus Funk (Employer: Cerence Operating Co)
- Tom Claes (Employer: Cerence Operating Co)
- Christophe Ris (Employer: Cerence Operating Co)
All named inventors were employed by Cerence Operating Co at the time of filing, with their interests in the patent assigned to the company. No unusual patterns, such as a mass departure of inventors, are discernible from the provided information.
Original assignee
The entity named on the issued patent, US12236947, is Cerence Operating Company. Cerence Inc. (which includes Cerence Operating Company) is an American multinational software company that develops artificial intelligence (AI) assistant technology primarily for the automotive and broader mobility and transportation markets. They provide a software platform for building conversational AI solutions, which are delivered on a white-label basis to automotive Original Equipment Manufacturers (OEMs). Their technology is directly embodied in products such as in-vehicle voice assistants, which leverage multimodal inputs including audio and visual cues to process commands, aligning directly with the claims of the '947 patent. Cerence Inc. was spun off from Nuance Communications in October 2019 and is currently an active, publicly traded company on Nasdaq under the ticker symbol CRNC.
Assignment timeline
The following assignment records are identified for US patent 12236947:
2023-07-10 (executed) / recorded 2023-07-10 — Reel 064199/0731
- Conveyance: Assignment
- Assignor: CLAES, TOM; QUAST, HOLGER; D'HOORE, BART; HALBOTH, CHRISTOPH; RIS, Christophe; SEPPI, DINO; FUNK, Markus
- Assignee: CERENCE OPERATING COMPANY
- Correspondent: Not explicitly available from Google Patents Legal Events.
- Context: Initial assignment of patent rights from the inventors to their employer, Cerence Operating Company, at the time of the patent application filing.
2024-04-12 (executed) / recorded 2024-04-15 — Reel 067417/0303
- Conveyance: Security Agreement
- Assignor: CERENCE OPERATING COMPANY
- Assignee: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
- Correspondent: Not explicitly available from Google Patents Legal Events.
- Context: Cerence Operating Company pledged the patent as collateral in a financing arrangement with Wells Fargo Bank.
2024-12-31 (executed) / recorded 2025-01-02 — Reel 069797/0422
- Conveyance: Release
- Assignor: WELLS FARGO BANK, NATIONAL ASSOCIATION
- Assignee: CERENCE OPERATING COMPANY
- Correspondent: Not explicitly available from Google Patents Legal Events.
- Context: Wells Fargo Bank released its security interest in the patent, restoring full unencumbered rights to Cerence Operating Company.
Note: Correspondent information for these records was not explicitly available in the provided Google Patents legal event data, and a live USPTO Assignment search cannot be performed by the model.
Timeline diagram
timeline
title Ownership of US 12236947
2023 : Inventors assigned to Cerence
2024 : Security agreement with Wells Fargo
2025 : Security interest released
: Patent issued to Cerence
NPE / troll-pattern signals
- Shell-entity transfer — Not present. The patent has consistently been held by Cerence Operating Company, an operating company with a clear business in automotive AI. The recorded transfers involve a security interest with a bank, not a transfer of ownership to a shell licensing entity.
- Known asserter in the chain — Not present. Neither Cerence Operating Company nor Wells Fargo Bank are identified as known patent assertion entities (NPEs) or patent trolls. Cerence is a product-shipping company, and Wells Fargo is a financial institution.
- Repeat correspondent across the chain — Unclear. Correspondent information is not available from the provided data for any of the recorded assignments (Reel 064199/0731, Reel 067417/0303, Reel 069797/0422). Therefore, it cannot be determined if a repeat correspondent is present.
- Cascading transfers — Not present. The recorded history shows an initial assignment, followed by a security agreement and its subsequent release. These are not multiple consecutive transfers through chained LLCs within a short timeframe.
- Pre-litigation transfer — Not present. As of the prior litigation summary (April 26, 2026), there is no known litigation involving US12236947. Thus, no transfers can be categorized as pre-litigation.
- Bankruptcy fire-sale — Not present. Cerence Operating Company is an active, publicly traded company, and there is no indication of bankruptcy.
- Privateering — Not present. The patent remains with Cerence Operating Company, an operating company. There is no evidence of a transfer to an NPE for assertion on Cerence's behalf.
- Defensive aggregator (anti-NPE) — Not present. The patent is currently held by Cerence Operating Company, not a defensive aggregator.
Verdict
Operating-company assertion
The patent US12236947 is currently owned by Cerence Operating Company, which is an active operating company developing and shipping products in the field of automotive AI that embody the patent's claims. The assignment records (Reel 064199/0731, Reel 067417/0303, Reel 069797/0422) reflect standard corporate intellectual property management and financing activities, not those typically associated with patent assertion entities.
For verification, please refer to the USPTO Assignment Center: https://assignmentcenter.uspto.gov/ and search for patent number 12236947.
Generated 5/29/2026, 9:04:22 PM
Prior art
Earlier patents, publications, and products that may anticipate or render the claims unpatentable.
Prior Art Analysis for U.S. Patent 12,236,947
An analysis of the prior art cited during the prosecution of U.S. Patent 12,236,947, "Flexible-format voice command," reveals several key references that the patent examiner considered. This review is critical in understanding the novel contributions of the '947 patent as determined by the United States Patent and Trademark Office (USPTO). The following analysis details the most relevant cited patents and their potential relationship to the claims of the '947 patent.
It is important to note that anticipation under 35 U.S.C. § 102 requires that a single prior art reference disclose each and every element of a claimed invention. The following analysis identifies claims that are potentially anticipated by the cited references, reflecting the examiner's likely considerations.
U.S. Patent Application Publication No. US2013/0297319A1
- Full Citation: Kim, Yongsin. "Mobile device having at least one microphone sensor and method for controlling the same." U.S. Patent Application Publication No. US2013/0297319A1, published November 7, 2013. Filed May 1, 2012.
- Brief Description: This patent application describes a mobile device that uses a microphone to detect a user's voice command. The system can activate and control functions based on the recognized voice command. It discusses using a voice trigger to initiate voice recognition.
- Potential Anticipation of Claims: This reference was likely considered in relation to the foundational aspects of voice command processing.
- Claim 1 & 18: The '319 application describes receiving an audio input and causing a system to act based on a command within that input. However, it does not appear to disclose the crucial element of simultaneously receiving and processing a video input to identify a visual characteristic of the user to determine their intent, which is a key limitation of claims 1 and 18 of the '947 patent.
- Claim 17: Similarly, while the '319 application deals with voice commands, it does not seem to incorporate the state of a dialog between the user and the system to determine if an utterance is a command.
U.S. Patent Application Publication No. US2019/0069017A1
- Full Citation: "Methods and systems for enhancing set-top box capabilities." U.S. Patent Application Publication No. US2019/0069017A1, published February 28, 2019. Filed August 31, 2017. Assignee: Rovi Guides, Inc.
- Brief Description: This application details methods for enhancing the functionality of a set-top box, including through voice commands. It discusses receiving user inputs, which can include voice, to control the device.
- Potential Anticipation of Claims: This reference addresses voice control in a specific consumer electronics context.
- Claim 1 & 18: The '017 application discloses receiving audio commands to control a system. However, like the '319 application, it does not appear to teach the combination of audio and video processing where a visual characteristic of the user is identified to help determine if an utterance is a command.
- Claim 17: The concept of using the state of a dialog is not a central feature of the '017 application's disclosure in the way it is claimed in the '947 patent.
U.S. Patent No. 10,388,272B1
- Full Citation: "Training speech recognition systems using word sequences." U.S. Patent No. 10,388,272B1, issued August 20, 2019. Filed December 4, 2018. Assignee: Sorenson IP Holdings, LLC.
- Brief Description: This patent focuses on the training of speech recognition systems. It describes methods for improving the accuracy of these systems by using specific word sequences during the training process.
- Potential Anticipation of Claims: This patent is relevant to the underlying speech recognition technology but less so to the specific multi-modal command determination method of the '947 patent. Its focus is on the training of the language model, which is a component of the system described in the '947 patent but not the core of the independent claims. It does not appear to disclose the claimed method of using combined audio and video inputs at the time of command issuance to determine user intent.
U.S. Patent Application Publication No. US2020/0310842A1
- Full Citation: "System for User Sentiment Tracking." U.S. Patent Application Publication No. US2020/0310842A1, published October 1, 2020. Filed March 27, 2019. Assignee: Electronic Arts Inc.
- Brief Description: This application describes a system for tracking user sentiment, which can involve analyzing voice data to determine a user's emotional state. This can be used to adapt a system's behavior.
- Potential Anticipation of Claims:
- Claim 1, 8, & 18: This reference is particularly relevant as it discusses analyzing user states, which could be interpreted to include visual cues. However, the '842 application's primary focus is on sentiment analysis rather than the specific problem of determining whether an utterance is a "command directed to a system" based on combined audio and video inputs, including identifying a "visual characteristic associated with the user uttering the first utterance" for this purpose. The '947 patent's claims are more specific about using this multi-modal analysis to solve the command-intent problem.
- Claim 17: The use of dialog state is not a primary teaching of this reference in the context of identifying a command.
U.S. Patent Application Publication No. US2020/0329297A1
- Full Citation: "Automated control of noise reduction or noise masking." U.S. Patent Application Publication No. US2020/0329297A1, published October 15, 2020. Filed April 12, 2019. Assignee: Bose Corporation.
- Brief Description: This application is focused on audio processing, specifically the control of noise reduction and masking technologies. It may involve analyzing audio to distinguish speech from background noise.
- Potential Anticipation of Claims: While this reference deals with advanced audio processing, it does not appear to disclose the multi-modal, intent-determining aspects of the '947 patent's independent claims. Its teachings are ancillary to the core invention claimed in the '947 patent.
In summary, while the cited prior art establishes a background for voice command and speech recognition technologies, none of the examined references appear to fully disclose the specific combination of features recited in the independent claims of U.S. Patent 12,236,947, particularly the use of video input to identify a user's visual characteristics in conjunction with audio processing to determine if an utterance is a system-directed command, and the further use of dialog context for this determination. This suggests that the novelty of the '947 patent, in the eyes of the examiner, resided in this specific multi-modal approach to discerning user intent.
Generated 5/8/2026, 10:06:53 PM
Obviousness
Combinations of prior art that suggest the claimed invention would have been obvious under 35 U.S.C. § 103.
Obviousness Analysis of US Patent 12,236,947
Date of Analysis: May 8, 2026
Patent under Review: US 12,236,947 ("the '947 patent")
Relevant Legal Standard: Under 35 U.S.C. § 103, a patent claim is invalid as obvious if the differences between the claimed invention and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art (POSITA).
This analysis examines the independent claims of the '947 patent in light of the cited prior art. The core of the invention, as outlined in the independent claims, is a multimodal system for recognizing voice commands by processing both audio and video inputs to determine user intent.
Claim 1 Analysis: Multimodal Command Recognition
Independent claim 1 claims a method of processing voice commands by:
- Receiving a first audio input.
- Receiving a first video input of the user.
- Determining the utterance is a system-directed command based on processing both the audio and video, including identifying a "visual characteristic" of the user.
- Causing the system to act on the command.
A potential obviousness argument could be constructed by combining the teachings of US 20130297319 A1 (Kim) and US 20200310842 A1 (Electronic Arts).
US 20130297319 A1 ("Kim") discloses a mobile device that uses a microphone and camera to recognize user commands. Kim's abstract and detailed description focus on activating voice recognition based on detecting a user's face and voice. This reference clearly establishes the use of both audio and video inputs for a voice-activated system.
US 20200310842 A1 ("Electronic Arts") describes a system for tracking user sentiment, which involves analyzing audio cues (like tone) and visual cues (like facial expressions) to gauge a user's emotional state. This system is designed to understand user intent and sentiment through multimodal analysis.
Motivation to Combine: A person of ordinary skill in the art developing voice-command systems would have been motivated to combine the teachings of Kim and Electronic Arts to improve the accuracy of command recognition. At the time of the invention, a well-known problem in voice-command systems was the "false trigger," where the system would incorrectly interpret ambient conversation as a command. A POSITA would recognize that simply detecting a face and a voice, as taught by Kim, is insufficient to confirm user intent. The sentiment and intent analysis taught by Electronic Arts, which explicitly uses facial expressions ("visual characteristics") and vocal tone, provides a direct solution to this problem. By integrating the intent-analysis methods of Electronic Arts with the basic multimodal command structure of Kim, a POSITA could create a more robust system that better distinguishes between a direct command and extraneous speech. This combination would lead to the system described in claim 1, as it would use both audio and visual characteristics to determine that an utterance is a system-directed command, thus rendering the claim obvious.
Claim 17 Analysis: Dialog State Context
Independent claim 17 builds upon claim 1 by adding the limitation of "using a state of a dialog between the system and the user in the determining." This means the system's understanding of the ongoing conversation influences whether an utterance is treated as a command.
This claim could be rendered obvious by combining Kim and Electronic Arts (as above) with the "Non-Patent Citations" referenced in the '947 patent's file history, specifically the 2013 IEEE paper by Wang et al., "Understanding computer-directed utterances in multi-user dialog systems."
- Wang et al. directly addresses the challenge of distinguishing system-directed speech from inter-user conversation in a multi-user environment. A key aspect of their research is using contextual cues, including the dialog history, to make this determination. The paper explicitly discusses how the system's state (e.g., whether it has just asked a question) is a critical factor in classifying a user's utterance.
Motivation to Combine: A POSITA, having already combined Kim and Electronic Arts to improve intent recognition, would naturally look to dialog context to further refine the system's accuracy. The problem of distinguishing commands from conversation is particularly acute in a continuous dialog. The Wang et al. paper provides a clear, well-documented method for using dialog state to solve this very problem. Therefore, a POSITA would have been motivated to incorporate the dialog-state analysis from Wang et al. into the multimodal framework of Kim and Electronic Arts. This would be a predictable improvement, allowing the system to understand, for example, that a short user utterance is likely a response to a system prompt rather than a new, unrelated command. This combination directly teaches all the elements of claim 17.
Claim 18 Analysis: The Physical System
Independent claim 18 claims the physical embodiment of the method: a system comprising an audio input device, a video input device, and a computing device configured to perform the multimodal command recognition.
This claim is rendered obvious by the same prior art that makes the method claims obvious.
- Kim discloses a system with the necessary hardware: a microphone (audio input device), a camera (video input device), and the mobile device's processor (a computing device).
- The combination of Kim, Electronic Arts, and Wang et al. teaches the functionality that the computing device would be configured to perform.
Since the underlying method is obvious, and the hardware components (microphone, camera, processor) are conventional and disclosed in the prior art for this exact purpose, the claim for the system itself would also have been obvious to a POSITA. There is no inventive concept in merely implementing an obvious method on a standard set of hardware components.
Conclusion
Based on the cited prior art, the independent claims of US patent 12,236,947 appear to be vulnerable to an obviousness challenge under 35 U.S.C. § 103. The combination of prior art references teaches the use of multimodal (audio and video) inputs to detect user commands, including the analysis of visual characteristics to determine intent, and the use of dialog context to improve accuracy. A person of ordinary skill in the art would have been motivated to combine these teachings to solve the well-known problem of false triggers and improve the overall user experience in voice-command systems.
Generated 5/8/2026, 10:06:52 PM
Extensions
Patent term adjustments, term extensions, continuations, divisionals, family members, and expiration dates.
Analysis of U.S. Patent 12,236,947
Date of Analysis: May 8, 2026
Based on the file history and data associated with U.S. Patent 12,236,947, the following is a detailed breakdown of its term, application history, and related family members.
Patent Term Adjustment (PTA) and Extensions (PTE)
- Patent Term Adjustment (PTA): A formal Patent Term Adjustment is calculated by the USPTO upon issuance to compensate for administrative delays during prosecution. This calculation is complex and depends on specific deadlines for office actions and applicant responses. The provided documentation for US 12,236,947 does not specify the exact number of PTA days granted. This information would typically be available on the front page of the issued patent or through the USPTO's Patent Center portal by reviewing the complete file wrapper. Without this specific data, a precise PTA calculation cannot be made.
- Patent Term Extension (PTE): There is no indication that this patent is eligible for or has been granted a Patent Term Extension. PTE is typically associated with delays in regulatory review for products like pharmaceuticals and is not applicable to this technology area.
Continuity and Family Data
- Continuation Applications: US Patent 12,236,947, which stems from application number US 18/219,906, is a continuation of a prior application. As stated in the patent's description, "This application is a continuation of U.S. application Ser. No. 17/239,894, filed on Apr. 26, 2021".
- Divisional Applications: The prosecution history does not indicate that any divisional applications have been filed from this patent's lineage.
- Patent Family Members: This patent is part of a family of two related U.S. patents. The relationship is as follows:
- Parent Application: U.S. Application No. 17/239,894, filed on April 26, 2021. This application subsequently issued as U.S. Patent No. 11,735,172 B2 on August 22, 2023.
- Child Application (Continuation): U.S. Application No. 18/219,906, filed on July 10, 2023. This is the application that matured into the patent under review, U.S. Patent No. 12,236,947 B2.
Projected Expiration Date
The term of a U.S. patent is twenty years from the filing date of the earliest U.S. non-provisional application to which it claims priority. In this case, the earliest effective filing date is that of the parent application (US 17/239,894), which is April 26, 2021.
Therefore, the base expiration date is calculated as:
April 26, 2021 + 20 Years = April 26, 2041.
This projected expiration date is also noted in the legal status information for the patent. It must be emphasized that this date does not include any Patent Term Adjustment (PTA) that may have been granted by the USPTO. Any awarded PTA would be added to this base expiration date, potentially extending the patent's term. The final expiration date can only be confirmed by accessing the official PTA calculation from the USPTO.
Generated 5/8/2026, 10:07:29 PM
Derivative works
Defensive disclosure: derivative variations of each claim designed to render future incremental improvements obvious or non-novel.
Defensive Disclosure and Prior Art Generation
Publication Date: May 8, 2026
Reference ID: DPD-T7-12236947
Title: System and Method for Multimodal, Context-Aware, and Failsafe Determination of Command Intent for Human-Machine Interfaces
This document discloses a series of technical implementations and derivative works related to the core concepts embodied in U.S. Patent 12,236,947. The purpose of this disclosure is to place these concepts into the public domain, thereby establishing prior art against future patent applications claiming these or similar incremental innovations.
Disclosures Pertaining to Claim 1: Multimodal Command Recognition
The core claim involves processing audio and video to identify a user's command intent. The following derivatives expand upon this concept.
Derivative 1.1: Component Substitution with Non-Visible Spectrum and Non-Acoustic Sensors
- Enabling Description: The system is implemented using a long-wave infrared (LWIR) thermal camera instead of a standard CMOS/RGB camera. The "visual characteristic" is a thermal signature change in the user's perioral and nasal regions, which corresponds to the exhalation patterns of speech. This method is effective in zero-light conditions. The audio input is supplemented by a bone conduction transducer pressed against the user's mastoid process, capturing vocal vibrations directly, rendering the system highly immune to ambient acoustic noise. Data from the LWIR sensor (as a 32x32 pixel thermal array) and the bone conduction sensor are fed into a convolutional neural network (CNN) for intent fusion.
- Mermaid Diagram:
graph TD A[User Utterance] --> B{LWIR Thermal Camera}; A --> C{Bone Conduction Transducer}; B --> D[Thermal Feature Extraction <br>(e.g., Perioral heat map)]; C --> E[Vibrational Audio Processing]; D & E --> F[Fusion CNN]; F --> G{Intent Classification <br>(Command / Non-Command)}; G --> H[System Action]; end
Derivative 1.2: Component Substitution with Gaze and Neurological Input
- Enabling Description: The video input device is a dedicated eye-tracking module using infrared emitters and sensors to calculate the user's precise point of gaze (POG). The "visual characteristic" is the user's gaze dwelling on a system-controllable object for >500ms concurrently with an utterance. This is combined with audio input and a consumer-grade electroencephalography (EEG) headband. The system identifies a command if the audio input is co-occurrent with a P300 event-related potential (ERP) signal from the EEG, indicating a recognition/decision event in the user's brain.
- Mermaid Diagram:
sequenceDiagram participant User participant EyeTracker participant EEG_Headband participant AudioMic participant FusionEngine User->>+AudioMic: Speaks "Activate that" User->>+EyeTracker: Looks at target device User->>+EEG_Headband: Brain registers decision (P300 wave) AudioMic->>FusionEngine: Provides audio stream EyeTracker->>FusionEngine: Provides Point of Gaze (POG) data EEG_Headband->>FusionEngine: Provides EEG data stream FusionEngine->>FusionEngine: Fuses POG, Audio, and P300 signal FusionEngine-->>User: Executes command on target device deactivate AudioMic deactivate EyeTracker deactivate EEG_Headband
Derivative 1.3: Cross-Domain Application in Sterile Surgical Environments
- Enabling Description: In a surgical robotics suite, a surgeon wears augmented reality (AR) glasses with an integrated microphone and an inward-facing camera for eye-tracking. When the surgeon issues a command like "increase power," the system determines intent. If the surgeon's gaze is directed at the robotic arm's control panel (displayed in the AR view), the command is routed to the robot. If their gaze is directed at a human nurse, the command is ignored by the robotic system and outputted via a speaker for the nurse. The "visual characteristic" is the POG relative to virtual objects in the AR overlay.
- Mermaid Diagram:
stateDiagram-v2 [*] --> Idle Idle --> Listening: Surgeon speaks Listening --> Intent_Analysis: Gaze Data Received state Intent_Analysis { [*] --> Gaze_On_Robot_UI: Eye tracking POG on AR robot controls Gaze_On_Robot_UI --> Route_To_Robot [*] --> Gaze_On_Human: Eye tracking POG on human colleague Gaze_On_Human --> Ignore_Command } Route_To_Robot --> Executed: Command sent to surgical robot Ignore_Command --> Idle: Command is for human staff Executed --> Idle
Derivative 1.4: Cross-Domain Application in Livestock Monitoring (AgTech)
- Enabling Description: An array of pan-tilt-zoom (PTZ) cameras and long-range microphones are installed in a cattle feedlot. The system continuously analyzes audio for bovine vocalizations indicative of distress (e.g., specific pitch and duration). Upon detecting such a vocalization, it directs the nearest camera to the source. The video stream is then analyzed to identify "visual characteristics" of distress, such as limping, isolation from the herd, or postural abnormalities. A "command" is determined if both audio and visual distress indicators are present, triggering an alert to the rancher's mobile device with the animal's tag number and location.
- Mermaid Diagram:
graph TD subgraph Monitoring_System A(Audio Analysis) -- Detects Distress Vocalization --> B(Cue Camera); B -- PTZ Camera focuses on source --> C(Video Analysis); C -- Identifies Visual Distress Signs --> D(Confirm Intent); end D -- Both Modalities Positive --> E{Send Alert}; E --> F[Rancher's Device]; C -- No Visual Distress --> A;
Derivative 1.5: Integration with AI for Predictive Intent Modeling
- Enabling Description: The system integrates a transformer-based neural network that is pre-trained on a massive dataset of human interactions. It receives the real-time audio and video streams (as sequences of feature vectors). Instead of just classifying the current utterance, the model predicts a probability distribution over a set of potential future commands the user might issue in the next 1-3 seconds. If a spoken command matches a high-probability prediction from the model, the confidence threshold for executing the command is lowered, allowing for faster, more responsive interaction, especially in high-noise environments where the audio signal may be degraded.
- Mermaid Diagram:
classDiagram class User { +Utterance +FacialData } class PredictiveIntentModel { -TransformerNetwork +process(audio_features, video_features) +predictNextCommands(top_k) : list } class CommandInterpreter { +transcribe(audio) : string +execute(command) } User --|> PredictiveIntentModel : provides features PredictiveIntentModel --|> CommandInterpreter : provides predictions User --|> CommandInterpreter : provides audio
Disclosures Pertaining to Claim 17: Dialog State Context
This claim adds the use of conversational context. The following derivatives expand this by redefining "dialog state" and its application.
Derivative 17.1: Expansion of "Dialog State" to Environmental and Biosensor State
- Enabling Description: The "state of a dialog" is expanded to include data from a network of IoT sensors. In a vehicle, this includes the current navigation route, weather conditions (from an external API), and cabin occupancy (from weight sensors in seats). The user's biometric state is monitored via a smartwatch, providing heart rate and galvanic skin response. An utterance like "it's getting dark" is interpreted as a command to turn on the headlights only if the IoT state confirms ambient light is below a set lumen threshold and the dialog state indicates no ongoing conversation about philosophy. An utterance like "I'm stressed" combined with high heart rate data will cause the system to suggest a calming playlist.
- Mermaid Diagram:
erDiagram USER { string utterance string biometric_state } SYSTEM { string dialog_history string environmental_state string vehicle_state } INTENT_PROCESSOR { string fused_context } USER ||--o{ INTENT_PROCESSOR : provides SYSTEM ||--o{ INTENT_PROCESSOR : provides
Derivative 17.2: Cross-Domain Application in Automated Educational Tutors
- Enabling Description: An AI-powered language tutor uses a comprehensive dialog state that tracks the student's learning history, including common grammatical errors and vocabulary weaknesses. When the student speaks a phrase, the system processes the audio. The video input is analyzed for facial cues of confusion (e.g., furrowed brow). If the student's spoken phrase contains a grammatical error previously flagged in the dialog state, and the visual cues indicate confusion, the system interrupts to provide a targeted correction. If no confusion is detected, it allows the conversation to flow, assuming a minor slip of the tongue. The "command" is an implicit request for help.
- Mermaid Diagram:
flowchart LR subgraph Student A[Speaks Phrase] B[Facial Expression] end subgraph TutorSystem C[Audio Processing] D[Video Processing] E[Access Dialog State <br> (Past Errors)] end A --> C B --> D F{Fuse Inputs & State} C & D & E --> F F -- Error matches past & Confusion detected --> G[Provide Correction] F -- Else --> H[Continue Conversation]
Derivative 17.3: The "Inverse" or Failure Mode: Stateless Privacy Mode
- Enabling Description: The system is designed with a user-selectable "stateless" mode. When activated, the system intentionally purges all dialog history after each interaction. It does not store logs of conversations or commands. In this mode, the determination of intent relies exclusively on the immediate audio and video input from the current utterance. This provides a higher degree of user privacy at the cost of contextual awareness. For example, the system cannot resolve anaphora (e.g., "turn it off") and will prompt for clarification, as it has no memory of what "it" refers to. This is a failsafe for privacy-sensitive applications.
- Mermaid Diagram:
stateDiagram-v2 state "Standard Mode" as S1 state "Stateless Mode" as S2 [*] --> S1 S1 --> S2: User Toggles Privacy S2 --> S1: User Toggles Privacy S1: Utterance -> Process(audio, video, history) -> Action S2: Utterance -> PurgeHistory -> Process(audio, video) -> Action/Prompt
Disclosures Pertaining to Claim 18: The Physical System
This claim covers the physical hardware. The following derivatives propose alternative and advanced hardware architectures.
Derivative 18.1: Distributed System Architecture using Edge Computing
- Enabling Description: The system is not a single computing device but a distributed network. The "audio input device" and "video input device" (e.g., cameras and mics in a smart home) are edge nodes. These nodes perform initial feature extraction locally using low-power processors (e.g., ARM Cortex-M series). The camera extracts facial landmark vectors, and the microphone extracts MFCCs (Mel-frequency cepstral coefficients). Only these low-bandwidth feature vectors are transmitted over the network to a central hub or cloud service for the final, computationally expensive intent fusion and command processing. This architecture preserves privacy (raw video/audio doesn't leave the room) and saves network bandwidth.
- Mermaid Diagram:
graph TD subgraph Edge_Device_1 A[Camera] --> B(Local Feature Extractor <br> Facial Landmarks); end subgraph Edge_Device_2 C[Microphone] --> D(Local Feature Extractor <br> MFCCs); end B -- Landmark Vector --> F{Central Fusion Hub}; D -- MFCC Vector --> F; F --> G[Intent Determination]; G --> H[Command Execution];
Derivative 18.2: Component Substitution with Neuromorphic Hardware
- Enabling Description: The "computing device" is a specialized neuromorphic processor (e.g., based on Loihi or SpiNNaker architecture). Both the audio and video sensor data are converted into asynchronous spike trains. The intent determination model is implemented as a Spiking Neural Network (SNN). This hardware architecture provides extreme low-power operation, making it suitable for always-on applications in battery-powered devices like wearables or AR glasses. The processing is event-driven, consuming power only when new visual or auditory information is detected.
- Mermaid Diagram:
sequenceDiagram participant Sensor_Video participant Sensor_Audio participant SpikingEncoder participant Neuromorphic_SNN participant Actuator Sensor_Video->>SpikingEncoder: Raw pixel data Sensor_Audio->>SpikingEncoder: Raw audio waveform SpikingEncoder->>Neuromorphic_SNN: Asynchronous Spike Trains Neuromorphic_SNN->>Neuromorphic_SNN: Processes spikes, determines intent Neuromorphic_SNN->>Actuator: Triggers command action
Combination Prior Art Scenarios with Open-Source Standards
Combination with OpenCV and WebRTC: An implementation is disclosed wherein the system operates entirely within a web browser. The video and audio input devices are a standard webcam and microphone accessed via the WebRTC
getUserMedia()API. The received video frames are processed client-side in a WebAssembly module that uses the OpenCV.js library to perform real-time facial landmark detection and head pose estimation. These "visual characteristics" are combined with the audio stream (which may be transcribed locally or sent to a server) to determine command intent, enabling any website to become a multimodal conversational agent without requiring plugins or dedicated hardware.Combination with Kaldi and MQTT: A system is disclosed for an industrial control environment (IIoT). Microphones distributed throughout a factory floor are the audio input devices. They run a lightweight version of the Kaldi speech recognition toolkit for keyword spotting. Cameras act as video input devices. When a worker speaks a potential command, the Kaldi spotter and the video feed (analyzed for gestures or gaze direction) provide inputs. The fused intent is then published as a message on a lightweight MQTT (Message Queuing Telemetry Transport) broker. Subscribed robotic arms or machinery act on the command, creating a robust, low-latency, and standards-based factory control system.
Combination with Android Open Source Project (AOSP) and TensorFlow Lite: A system is disclosed as a modification to the core AOSP accessibility services. The audio input is from the device microphone and the video input is from the front-facing camera. A TensorFlow Lite model, optimized for mobile NPUs, is integrated into the OS. This model continuously processes the audio/video streams to determine if a user with motor impairments is attempting to issue a command versus speaking to someone else in the room. This multimodal intent signal is then used to activate the standard AOSP Voice Access service, reducing false activations and making the device more usable.
Generated 5/8/2026, 10:08:20 PM
Keep exploring
Other patents in High-Tech (T)
- US 10576716Here is a concise summary of US patent 10576716: Patent Number: US10576716B2 Title: Protective element and method for manufacturing display device Current Assignee: Magnolia White Corp (as of July 22, 2025) Original Assignee: Japan Display…
- US 12313913US patent 12313913, titled "System for powering head-worn personal electronic apparatus," was filed on March 6, 2024, and granted on May 27, 2025. The patent is assigned to Ingeniospec LLC, with Thomas A. Howell, David Chao, C. Douglass…
- US 9991030Here's a concise summary of US Patent 9991030: US Patent 9991030: High Performance Data Communications Cable Title: High performance data communications cable Assignee: Belden Inc. Inventors: Andrew John Wehrli, William Thomas Clark, Galen…
- US 8836842US Patent 8836842, titled "Capture mode outward facing modes," is currently active and set to expire on November 6, 2032. Here's a concise summary of the patent: Title: Capture mode outward facing modes Assignee: Multifold International…
- US 10482293Here's a concise summary of US patent 10482293: Patent Number: US104822293B2 Title: Interrogator and interrogation system employing the same Current Assignee: Lone Star SCM Systems LP Original Assignee: Medical IP Holdings LP Inventors…
- US 8139544Here is a concise summary of US patent 8139544: Title: Pilot tone processing systems and methods Assignee: Integral Wireless Technologies LLC (Previously assigned to Intellectual Ventures I LLC, Intellectual Ventures Assets 199 LLC, among…
- US 7738595Here is a concise summary of US patent 7738595: US Patent 7738595: Multiple input, multiple output communications systems Title: Multiple input, multiple output communications systems Assignee: Integral Wireless Technologies LLC Inventor…
- US 7676007Here's a concise summary of US Patent 7676007: US Patent 7676007 Summary Title: System and method for interpolation based transmit beamforming for MIMO-OFDM with partial feedback Current Assignee: Integral Wireless Technologies LLC…